Voidly’s multi-source Bayesian corroboration classifier
(corroboration_v1) was reported at
ROC AUC 0.92. It fuses four sensor networks —
OONI, IODA, CensoredPlanet and the Voidly probe network — into a
single naive-Bayes posterior, and it feeds the
auto-incident-watchdog as the “does an independent
source agree?” gate. A platform-wide audit this week found
several Voidly models reporting inflated metrics from shuffled
train/test splits that leak temporal autocorrelation — the 7-day
forecast’s “0.954 AUC” was really ~0.33 on the days
that matter. The corroboration model had never been individually
audited. This finding is that audit.
The first thing the audit checked was the split. Unlike the leaky
forecast, train-bayesian-corroboration.py uses a
forward-temporal split: train on the oldest 60 days,
test on the strictly-future newest 30 days, never shuffled. We
reproduced its number exactly — temporal test AUC
0.9157 — and a leave-country-out (LOCO)
cross-validation also held up at median AUC 0.866. So the leakage is
not autocorrelated country-day rows bleeding across folds.
The leakage is circular features. The model’s
label, is_censorship, is taken from the incidents table.
The model’s features, X_present, ask “did
source X emit an anomalous-level evidence row on this country-day?”
Those are the same evidence rows. An incident in Voidly is
minted from anomalous evidence: a SQL join through the
incident_evidence table shows that
343 of 344 confirmed censorship/mixed incidents have a
linked elevated/warning/critical
evidence row on the exact same country-day as
first_seen. The feature does not predict the label —
it partially is the label, re-derived.
If the “multi-source fusion” were doing real work, it would clearly beat any single source. It does not. The breakdown of which source minted each incident tells the story:
| How the 343 censorship/mixed incidents were sourced | Count |
|---|---|
censoredplanet only | 265 |
ooni-historical-backfill only | 56 |
ooni only | 35 |
ioda + censoredplanet | 43 |
ooni + censoredplanet | 1 |
77% of all positive labels were created from CensoredPlanet evidence
alone. So censoredplanet_present — a single raw
binary feature, no model at all — is almost a copy of
is_censorship. Scored directly:
| “Model” | Temporal test AUC |
|---|---|
| Full 4-source naive-Bayes corroboration | 0.9157 |
censoredplanet_present alone (1 raw feature) | 0.8997 |
ooni_present alone (1 raw feature) | 0.7069 |
The four-source Bayesian fusion adds +1.6pp AUC over a single circular feature. A 2,000-sample bootstrap on that lift gives a 95% confidence interval of [0.0pp, 3.3pp] — the interval all but touches zero (P(lift ≤ 0) = 0.024). The “multi-source corroboration” is, in effect, CensoredPlanet’s labeling trace with three near-inert extra terms. IODA actually carries a negative likelihood ratio for presence (LR 0.29) because IODA disruption rows are deliberately excluded from the censorship label; the Voidly-probe term fires on <1% of positives.
The metrics file reports F1, precision and recall at threshold 0.5 as
0.0, 0.0, 0.0. The posterior never crosses 0.5, even
on true positives, because the empirical prior is ~3.5% and the
likelihood ratios cannot pull a 3.5% prior past 50% on the source
combinations that actually occur. corroboration_v1 is a
ranker of “did CensoredPlanet flag this day”, not a
usable classifier. AUC — a rank metric — is the only thing
that looked good, and AUC is exactly the metric the circular feature
inflates.
To measure real skill we need a target the same-day features cannot encode. We re-ran the identical naive-Bayes pipeline on a genuine forecasting target: predict NEXT-day censorship from TODAY’s source presence, forward-temporal split. The AUC falls to 0.7348. And even that 0.73 is generous — a CensoredPlanet block today frequently recurs tomorrow and mints another CP-sourced incident, so persistence still leaks in. The honest read is that once the feature can no longer copy the label, the model’s discrimination collapses from 0.92 toward 0.73, and the residual is mostly autocorrelation rather than corroboration.
The headline AUC 0.92 is real arithmetic but near-tautological, and it is not a measure of censorship-detection skill. There is no model change to promote — the audit’s job was to find out whether the 0.92 was honest, and it is not. We did not retrain or swap the model, because any “improvement” measured against a circular label would be just as fake as the original 0.92. The honest fix is disclosure plus a data-pipeline change, not a new classifier:
corroboration_v1_metrics.json)
now carries a full leakage_audit block and seven
rewritten honest_caveats; promoted is set
to false. The live
/v1/classifier/corroborate/info endpoint serves them
verbatim.model_changelog.json, surfaced
at /atlas/changelog) gains a
corroboration-v1-bayesian-reeval entry recording the
reported 0.92, the reproduced 0.9157, the trivial-baseline 0.8997,
the +1.6pp lift, and the leakage-free 0.7348.auto-incident-watchdog.py docstring is corrected
(it previously claimed “AUC ~0.78”). The watchdog gates
draft incidents on posterior ≥ 0.5; because the
posterior almost never reaches 0.5, that gate is in practice a
conservative near-veto — it suppresses drafts, which
is safe, but it is NOT the independent confirmation it was
presented as. Two correlated signals derived from the same evidence
are not corroboration.The fix is not modeling, it is decoupling the feature from the label. Genuine corroboration needs at least one of:
Negative results count. The corroboration model still has a narrow legitimate use — as a conservative suppressor inside the editorial-draft watchdog — but its 0.92 was never evidence that Voidly’s four sensor networks independently agree. Publishing that plainly, next to the promoted experiments, is what keeps the Atlas methodology honest.
Audit script:
scripts/audit-bayesian-corroboration.py
(deterministic, seed 42; run on the Vultr ML server as the service
user). It reproduces the temporal AUC, the LOCO sweep, the
trivial-baseline comparison, the bootstrap CI on the lift, the
leakage-free next-day AUC, and the incident_evidence
circularity join, then writes the leakage_audit block into
/opt/voidly-ai/ml-deploy/corroboration_v1_metrics.json
(a .bak.preaudit copy is kept). Model unchanged:
corroboration_v1.pkl. Feature panel:
scripts/build-corroboration-features.py.
Training script (forward-temporal split, unchanged):
scripts/train-bayesian-corroboration.py.