Evidence, compared
Multi-source Bayesian corroboration
Source evidence combined for each country-day. These model probabilities depend on calibration and source coverage.
Trained 2026-05-21 · 183 confirmed censorship incidents · Raw JSON · Model info
Top country-days, last 30 days (posterior > 0.2)
Reading this view
One probability per country-day: given what OONI, IODA, CensoredPlanet, and Voidly probes observed, what is the chance this is real censorship? Naive-Bayes fusion with empirical likelihoods. Resolves journalist's question: “is this just one source's false positive?”
Per-source likelihoods
How often each source “fires” on labeled censorship days vs background days. LR present is the likelihood ratio when the source signals; values above 1 push the posterior up, below 1 push it down.
| Source | P(present | censorship) | P(present | not) | LR present | Δ AUC if removed |
|---|---|---|---|---|
| OONI | 54.1% | 30.7% | 1.76 | +1.1pp |
| IODA | 14.8% | 51.3% | 0.29 | -0.4pp |
| CensoredPlanet | 99.2% | 18.8% | 5.29 | +23.0pp |
| Voidly probes | 0.8% | 0.0% | 27.34 | -0.0pp |
Scroll to see every column.
Methodology
Methodology
Each country-day is one observation. Per source s, we compute the presence indicator: did s emit any elevated/warning/critical-level signal on that day? We then estimate two likelihoods on the training window:
P(s present | C=1): how oftensfires on labeled censorship daysP(s present | C=0): how oftensfires on background days
We use Laplace smoothing (α=1) on both branches so no source produces a zero or infinite likelihood. The posterior is computed in log-odds space for numerical stability:
log_odds(C=1) = log(prior/(1-prior))
+ Σ_s log(LR(s = observed))
LR(s=present) = P(s present | C=1) / P(s present | C=0)
LR(s=absent) = P(s absent | C=1) / P(s absent | C=0)
posterior = sigmoid(log_odds)Training window: 2026-02-20 to 2026-04-21. Held-out test: 2026-04-21 to 2026-05-21(63 positives, 2,485 rows total).
Honest caveats
- LEAKAGE AUDIT (2026-05-22): the reported AUC ~0.92 is real arithmetic but near-tautological. The label is_censorship is derived from the incidents table; 343 of 344 censorship/mixed incidents were minted from an anomalous evidence row on the same country-day the feature counts. The feature partially ENCODES the label.
- censoredplanet_present alone (one raw binary feature, no model) scores AUC ~0.90 on the temporal test set. The 4-source Naive Bayes adds only ~1.6pp (bootstrap 95% CI [0.0, 3.3pp]) — the 'multi-source fusion' contributes essentially nothing.
- On a leakage-free target (predict NEXT-day censorship from TODAY's sources) AUC falls to ~0.73.
- At threshold 0.5 the model's F1/precision/recall are all 0.0 — it is a ranker, not a classifier. Treat the posterior as a relative score only.
- The split itself is honest (forward-temporal, not shuffled). The leakage is circular features, not split contamination — but the effect on the headline number is the same: the metric overstates real skill.
- Naive Bayes assumes per-source independence given the label, violated when sources cover overlapping ASNs.
- The corroboration gate in auto-incident-watchdog.py is a conservative near-veto, not the independent confirmation it appears to be.