Voidly Atlas runs two production models: the v3.3 censorship classifier (detects censorship in already-collected country-day evidence) and the v1 7-day forecast (predicts shutdown risk). Both are retrained on a weekly cron, and a dual-holdout gate blocks a regressing model from being promoted. But there was no principled detector for the question that sits upstream of the gate: has the live data distribution drifted away from what the models were trained on?
That gap had a concrete cost. In March–April 2026 a labeling
bug let raw IODA connectivity-disruption alerts (fiber cuts, BGP
leaks, weather outages) flow into the forecast’s
target_7day labels as if they were confirmed
censorship. The April positive rate exploded to 79%
against a training baseline near 5%. The model dutifully learned to
call almost everything a shutdown. The bug festered for
weeks before a human noticed. A distribution-shift
detector watching the label rate would have flagged it on day 2.
The concept-drift detector
(scripts/detect-concept-drift.py, daily 05:20 UTC)
closes that gap. It is a monitor, not a gate — it does not
block anything; it raises a verdict and, when the verdict is severe
enough, queues an early retrain.
A one-time baseline builder
(scripts/build-concept-drift-baseline.py, also re-run
after every retrain) freezes the training-time
distribution of each model’s features: mean, standard
deviation, a 7-point quantile summary, and — critically
— the ten decile bin edges of the training
distribution plus the training proportion in each bin. Freezing the
bins is what makes the daily comparison honest: live data is binned
into the exact same edges.
Every day the detector recomputes each model’s features on the last 7 days of live data, using the same derivation code path as the training pipeline (the classifier’s country-day aggregation and the forecast’s feature parquet), and runs two tests per feature:
sum((live% − train%) × ln(live% / train%))
over the ten frozen bins. PSI is the primary signal. Standard
industry thresholds: > 0.2 = significant
drift, > 0.25 = major drift.
Separately it tracks label drift: the positive rate of the target over the trailing 30 days versus the training baseline. This is the metric that would have caught the IODA disaster.
A composite drift_score in [0, 1] combines 60% mean
feature PSI and 40% label drift, and is floored at
retrain-recommended whenever label drift alone is “major”
— so the IODA failure mode always triggers a retrain
regardless of feature PSI. The score maps to a verdict:
stable / watch /
drifted / retrain-recommended.
When a model crosses retrain-recommended, the
detector writes a flag to the existing drift-trigger queue
(/opt/voidly-ai/data/retrain-queue.json), which
weekly-retrain.sh already reads — so the next
retrain runs early. It deliberately does not
auto-retrain on the spot: retrains are expensive, and a 12-hour
cooldown serializes drift-driven retrains with the weekly cron.
A naive feature-by-feature PSI monitor has a built-in bug, and the
first run surfaced it loudly. The features month,
week_of_year and day_of_week posted PSI
values of 13 to 17 — orders of magnitude past
any threshold. That is not drift. A 7-day live window only ever
spans one calendar month and one or two ISO weeks, so its decile
distribution over a calendar feature can never match a
baseline drawn from a full year. The PSI is structurally guaranteed
to be enormous and says nothing about model health.
The detector flags those five calendar features
(month, week_of_year,
day_of_week, is_weekend,
is_friday) as cyclical, still reports their
PSI for transparency, but excludes them from the composite
drift_score. Before that fix both models read
retrain-recommended on calendar noise alone. After it, the verdicts
are honest.
On 21 May 2026 the detector scored both models:
block_rate lag and rolling-mean features — 14
of the 38 forecast features — show PSI around 3.4. Over the
identical 21-country set, the mean 7-day block rate jumped roughly
tenfold (0.023 in the historical baseline to 0.234
in the trailing week). The forecast’s recent-history inputs
are sitting well outside the distribution the current model
learned from. Forecast label drift is a modest 0.10 — not
the trigger; the feature drift is.
The detector queued the early retrain via the shared queue (subject to the 12-hour cooldown, which deferred to an already-pending entry on the first run).
neighbor_*) are not monitored: they need the offline
adjacency and regime-correlation pipeline and cannot be recomputed
on a 7-day live window.
Live at GET /v1/atlas/concept-drift (both models,
per-feature PSI/KS, drift_score, verdict, label drift),
GET /v1/atlas/concept-drift/{model} for one model
(classifier-v3.3 or forecast-v1; short
aliases classifier / forecast accepted),
and GET /v1/atlas/concept-drift/info for the full
methodology. Refreshed daily at 05:20 UTC. Implementation:
scripts/build-concept-drift-baseline.py +
scripts/detect-concept-drift.py +
scripts/patch-concept-drift-endpoint.py.