Voidly Atlas ships a shutdown-duration model: a Random Survival
Forest (RSF) that answers “once an internet shutdown or censorship
event starts, how long will it last?” The model card reported a
concordance index (c-index) of 0.728 — comfortably
above the 0.65 promote floor, flagged passed_promote_floor: true,
and served live at POST /v1/forecast/duration. This finding is
an individual honest audit of that number, prompted by a platform-wide ML
review that caught several Voidly models reporting inflated metrics from
shuffled train/test splits. The duration RSF had not yet been audited. It
has now, and the headline number does not survive.
The training script scripts/train-shutdown-duration-rsf.py
evaluates the RSF with a single call to scikit-learn’s
train_test_split(…, test_size=0.25, random_state=42) —
a random 75/25 shuffle, stratified only on the censoring
flag. There is no temporal ordering and no country grouping. The
c-index 0.728 is the concordance of the RSF’s risk scores on that one
random test fold of 86 incidents.
Three problems make that number unreliable, in increasing order of severity.
Of 343 incidents, only 74 have an observed end (the rest are
right-censored: 78.4% censoring). A 25% test fold therefore contains roughly
19 events, and a c-index is computed from event-vs-other pairs — the
reported run had just 209 concordant + 77 discordant pairs. That is a tiny,
noisy sample. Re-running the identical random split across 20 different
seeds gives a c-index of 0.728 ± 0.055, ranging from
0.62 to 0.82. The published 0.728 is simply where random_state=42
happened to land in a wide band — not a stable estimate of skill.
The single most informative feature, by permutation importance, is
first_seen_year (and the runner-up,
prior_shutdowns_count_24mo, is also a recency proxy). Here is
why that is leakage rather than signal: the “event observed”
label is defined as status = 'confirmed', and
confirmed status is overwhelmingly an artifact of incident
age, not of the shutdown actually ending.
| Incident year | Incidents | Observed-end rate |
|---|---|---|
| 2019–2023 | 27 | ~100% |
| 2024 | 33 | 15% |
| 2025 | 25 | 56% |
| 2026 | 253 | 9% |
Old incidents have been curated and marked confirmed; recent
ones are still active. A random shuffle puts old and new years
in both train and test, so the model learns “old year ⇒ this row
is one of the resolved ones” — a pattern that exists only because
of how the dataset was assembled, and that cannot generalise to a live
shutdown happening today.
The correct evaluation for a duration model is a forward-temporal
split: train on shutdowns whose outcome was known before a cutoff,
test on strictly later ones. Sorting incidents by their end date
(last_seen) and sweeping the cutoff year:
| Train cutoff | Honest temporal c-index |
|---|---|
| ≤ 2022 | 0.609 |
| ≤ 2023 | 0.563 |
| ≤ 2024 | 0.571 |
| ≤ 2025 | 0.437 |
The mean honest forward-temporal c-index is about 0.55, and the most recent fold — the one closest to how the model is actually used — scores 0.44, below a coin flip. A leave-country-out cross-validation pooled across all 65 countries lands at 0.71, but LOCO still mixes early and late incidents within each fold, so it inherits the same year leak; it is not the honest number for a model whose job is to predict the future.
The promote gate requires a duration model to beat a naive baseline —
predict this country’s mean past observed duration. On the identical
forward-temporal folds, that naive baseline scores a mean c-index of
0.509. The RSF’s honest 0.55 is within noise of it, and
on two of the four folds the RSF scores below naive. We also gave
the model its best fair shot: an enriched feature set adding incident
confidence, measurement count, anomaly rate, source count, affected-domain
/service/ASN counts, severity grade and blocking mechanism — and
dropping the leaky first_seen_year. The enriched RSF scored
0.495 on the forward-temporal split. More features did not
help. There is no version of this model that clears the gate.
| Evaluation | c-index | Honest? |
|---|---|---|
| Published (random split, seed 42) | 0.728 | no — leaks year |
| Random split, mean of 20 seeds | 0.728 ± 0.055 | no — still leaks year |
| Leave-country-out, pooled | 0.712 | partial — still mixes time |
| Forward-temporal, base features | 0.559 | yes |
| Forward-temporal, enriched features | 0.495 | yes |
| Naive baseline (country mean duration) | 0.509 | yes |
| Coin flip | 0.500 | — |
No duration model — the original RSF or the enriched rebuild — beats a naive country-mean baseline on a leak-free forward-temporal split. Per the promote gate, nothing is promoted. The honest finding is that, with the data Voidly currently holds, how long a shutdown will last is not predictable from country, region, severity, prior-shutdown history or event context. The reasons are structural and worth stating plainly:
status = 'confirmed' tracks how long ago the incident was
filed, not when blocking stopped. Until end dates come from genuine
measurement (the day probes stop seeing the block), the survival target is
contaminated.last_seen − first_seen
rounded to ingestion cadence, not precise shutdown lengths.
This does not mean the question is unanswerable forever — it means it
is not answerable with today’s data, and claiming a 0.728 c-index
implied otherwise. The model artifact and its sidecar metrics have been
corrected to carry the honest forward-temporal c-index (~0.50–0.56),
passed_promote_floor: false, and an explicit
honest_caveats block; the /v1/forecast/duration
endpoint now surfaces the audit verdict in every response. Publishing the
honest negative — and not shipping a number we cannot stand behind
— is the fix.