Models rot silently when production data drifts away from the training distribution. DriftWatch Pro registers a training baseline per feature, then scores each live batch with two complementary tests and flags drift before accuracy quietly degrades:
- Two-sample Kolmogorov–Smirnov (KS) — detects a change in distribution shape (with an asymptotic p-value).
- Population Stability Index (PSI) — the standard MLOps drift metric on binned frequencies.
Both are implemented directly on numpy (no scipy), so the maths is auditable and the package is light.
MLOps reliability, local-first — the statistics run locally, cheaply, and explainably.
git clone https://github.com/Kimosabey/driftwatch-pro.git
cd driftwatch-pro
pip install -r requirements.txt
python -m unittest discover -s tests # 10 tests
python -m driftwatch.demo # baseline vs. stable / shifted / variance batches
docker compose up # HTTP API on :8000curl -s localhost:8000/baseline -H 'content-type: application/json' \
-d '{"feature":"latency","values":[ ...training sample... ]}'
curl -s localhost:8000/check -H 'content-type: application/json' \
-d '{"feature":"latency","values":[ ...live batch... ]}'
# -> {"feature":"latency","ks_stat":0.33,"p_value":0.0,"psi":0.65,"drifted":true,"severity":"high"}batch KS D p-value PSI verdict
stable 0.0452 0.0645 0.0091 ok
mean-shift 0.3328 0.0000 0.6474 DRIFT (high)
variance 0.1372 0.0000 0.3340 DRIFT (high)
- Two detectors — KS-test (shape) + PSI (binned stability), so shifts one test misses the other catches.
- Per-feature baselines with a simple register-then-check API.
- Severity grading —
none/moderate/highfrom p-value and PSI thresholds.
- Adaptive binning — PSI bin count scales to sample size, so small batches don't false-positive.
- Dependency-light — numpy only; tests run on the standard-library test runner.
%%{init: {'theme':'base','themeVariables':{'primaryColor':'#ffffff','lineColor':'#6366F1','mainBkg':'#ffffff'}}}%%
graph LR
A([Training data]) --> B([Baseline per feature])
C([Live batch]) --> D{Detectors}
B --> D
D --> E([KS-test: shape])
D --> F([PSI: stability])
E --> G([Severity + alert])
F --> G
style B fill:#eef2ff,stroke:#6366F1,stroke-width:2px,color:#3730a3
style D fill:#e0e7ff,stroke:#6366F1,stroke-width:2px,color:#3730a3
style G fill:#eef2ff,stroke:#6366F1,stroke-width:2px,color:#3730a3
The hard part is telling real drift from sampling noise: choosing tests and thresholds (and adapting PSI bins to sample size) so alerts fire on signal, not variance. See docs/ARCHITECTURE.md.
| Layer | Technology | Role |
|---|---|---|
| Language | Python 3.12 | Detectors + monitor + API |
| Numerics | numpy | KS statistic, PSI, quantile binning |
| Transport | stdlib http.server |
Zero-framework HTTP API |
| Tests | stdlib unittest |
10 deterministic tests |
| Container | Docker + Compose | One-command run |
- Architecture — detectors, thresholds, adaptive binning
- Getting Started · Failure Scenarios · Interview Q&A
- Categorical-feature drift (chi-square)
- Rolling-window monitoring + alert history
- Auto-retrain trigger when severity stays high
- Grafana / Prometheus metrics export
Released under the MIT License.
Harshan Aiyappa Senior Full-Stack Hybrid AI Engineer Voice AI • Distributed Systems • Infrastructure



