Voidly Atlas ships a shutdown-duration model: a Random Survival Forest (RSF) that answers “once an internet shutdown or censorship event starts, how long will it last?” The model card reported a concordance index (c-index) of 0.728 — comfortably above the 0.65 promote floor, flagged passed_promote_floor: true, and served live at POST /v1/forecast/duration. This finding is an individual honest audit of that number, prompted by a platform-wide ML review that caught several Voidly models reporting inflated metrics from shuffled train/test splits. The duration RSF had not yet been audited. It has now, and the headline number does not survive.

How the 0.728 was computed

The training script scripts/train-shutdown-duration-rsf.py evaluates the RSF with a single call to scikit-learn’s train_test_split(…, test_size=0.25, random_state=42) — a random 75/25 shuffle, stratified only on the censoring flag. There is no temporal ordering and no country grouping. The c-index 0.728 is the concordance of the RSF’s risk scores on that one random test fold of 86 incidents.

Three problems make that number unreliable, in increasing order of severity.

Problem 1 — the point estimate is a lucky seed

Of 343 incidents, only 74 have an observed end (the rest are right-censored: 78.4% censoring). A 25% test fold therefore contains roughly 19 events, and a c-index is computed from event-vs-other pairs — the reported run had just 209 concordant + 77 discordant pairs. That is a tiny, noisy sample. Re-running the identical random split across 20 different seeds gives a c-index of 0.728 ± 0.055, ranging from 0.62 to 0.82. The published 0.728 is simply where random_state=42 happened to land in a wide band — not a stable estimate of skill.

Problem 2 — the random split leaks calendar time

The single most informative feature, by permutation importance, is first_seen_year (and the runner-up, prior_shutdowns_count_24mo, is also a recency proxy). Here is why that is leakage rather than signal: the “event observed” label is defined as status = 'confirmed', and confirmed status is overwhelmingly an artifact of incident age, not of the shutdown actually ending.

Incident yearIncidentsObserved-end rate
2019–202327~100%
20243315%
20252556%
20262539%

Old incidents have been curated and marked confirmed; recent ones are still active. A random shuffle puts old and new years in both train and test, so the model learns “old year ⇒ this row is one of the resolved ones” — a pattern that exists only because of how the dataset was assembled, and that cannot generalise to a live shutdown happening today.

Problem 3 — the honest split shows no skill

The correct evaluation for a duration model is a forward-temporal split: train on shutdowns whose outcome was known before a cutoff, test on strictly later ones. Sorting incidents by their end date (last_seen) and sweeping the cutoff year:

Train cutoffHonest temporal c-index
≤ 20220.609
≤ 20230.563
≤ 20240.571
≤ 20250.437

The mean honest forward-temporal c-index is about 0.55, and the most recent fold — the one closest to how the model is actually used — scores 0.44, below a coin flip. A leave-country-out cross-validation pooled across all 65 countries lands at 0.71, but LOCO still mixes early and late incidents within each fold, so it inherits the same year leak; it is not the honest number for a model whose job is to predict the future.

Does it beat a naive baseline? No.

The promote gate requires a duration model to beat a naive baseline — predict this country’s mean past observed duration. On the identical forward-temporal folds, that naive baseline scores a mean c-index of 0.509. The RSF’s honest 0.55 is within noise of it, and on two of the four folds the RSF scores below naive. We also gave the model its best fair shot: an enriched feature set adding incident confidence, measurement count, anomaly rate, source count, affected-domain /service/ASN counts, severity grade and blocking mechanism — and dropping the leaky first_seen_year. The enriched RSF scored 0.495 on the forward-temporal split. More features did not help. There is no version of this model that clears the gate.

Evaluationc-indexHonest?
Published (random split, seed 42)0.728no — leaks year
Random split, mean of 20 seeds0.728 ± 0.055no — still leaks year
Leave-country-out, pooled0.712partial — still mixes time
Forward-temporal, base features0.559yes
Forward-temporal, enriched features0.495yes
Naive baseline (country mean duration)0.509yes
Coin flip0.500—

Verdict: honest negative — shutdown duration is not predictable here

No duration model — the original RSF or the enriched rebuild — beats a naive country-mean baseline on a leak-free forward-temporal split. Per the promote gate, nothing is promoted. The honest finding is that, with the data Voidly currently holds, how long a shutdown will last is not predictable from country, region, severity, prior-shutdown history or event context. The reasons are structural and worth stating plainly:

This does not mean the question is unanswerable forever — it means it is not answerable with today’s data, and claiming a 0.728 c-index implied otherwise. The model artifact and its sidecar metrics have been corrected to carry the honest forward-temporal c-index (~0.50–0.56), passed_promote_floor: false, and an explicit honest_caveats block; the /v1/forecast/duration endpoint now surfaces the audit verdict in every response. Publishing the honest negative — and not shipping a number we cannot stand behind — is the fix.