Voidly Atlas runs eight OONI test types against probe targets every
six hours: web_connectivity, signal,
whatsapp, telegram, facebook_messenger,
tor, http_invalid_request_line, and
http_header_field_manipulation. Until today we treated all
eight as equal contributors to country-day censorship labels. The
OONI test-type meta-classifier asks a different question: for
each country, which test type carries the strongest diagnostic
signal of actual censorship?
Why this matters
Different censors leave different fingerprints. Iran shoots Tor
bridges in the obfs4 handshake — so the tor test type
separates Iranian incident-days from quiet days at AUC 0.84,
better than any other test. Russia interferes broadly with web
protocols — web_connectivity gives RU a striking
AUC 0.98. When an analyst opens a country page and sees a
candidate event, this ranking tells them where to look first:
load the tor explorer for IR, the web_connectivity explorer for
RU. It is a UX primitive for the OONI Explorer.
Methodology
For every country with at least 50 labelled days of OONI data we
fit a small per-test-type model:
- Bucket evidence at (country, day, test_type),
extracting test_type from the OONI Explorer URL embedded in
each evidence row's
source_url.
- Compute anomaly rate per bucket. Where OONI
preserves the “N/M measurements anomalous”
ratio in
upstream_claim, we use that exact rate;
otherwise we fall back to the fraction of non-ok
signal_type rows.
- Label days: a (country, day) is positive iff
a confirmed censorship or mixed incident's
[first_seen, last_seen] window (± 1 day) covers
that day. We use the same 343-incident label set the v3.3
classifier and the multi-horizon forecast train on.
- Fit logistic regression (sklearn,
class_weight='balanced', standardized input)
per (country, test_type) pair, requiring ≥5 positive and ≥5
negative days. Score AUC in-sample. The signed coefficient
captures whether a high anomaly rate raises or
lowers the predicted censorship probability — useful
when a test type sees rare false positives that flip the sign.
- Rank the eight test types per country. The
top_test_type is the field a journalist should query
first; the auc_range (max minus min across viable
types) tells you how much it actually matters.
Results
30 countries cleared the 50-labelled-day floor.
13 of them show an AUC range > 0.10 between
their best and worst test type — i.e. the choice of which OONI
probe to run actually matters. Globally,
web_connectivity wins #1 most often (12 countries),
followed by tor (5), http_invalid_request_line
(4), and facebook_messenger, signal,
whatsapp, http_header_field_manipulation at
2 each. The exact distribution is published in the
top_test_type_distribution field of the
/info endpoint.
Spotlight countries
- Iran (IR): top test type is tor
at AUC 0.838, with 71 days and 22 incident-days. The
next-best (http_header_field_manipulation) trails
by 0.107 AUC. The negative coefficient (-1.18) confirms the
classic Iranian pattern: when tor anomaly rate spikes, an
incident is happening.
- Russia (RU): web_connectivity
dominates at AUC 0.977 with 622 incident-days and 8
viable test types. The AUC range is 0.468 — the widest of
any country — driven by TSPU's broad-protocol interference
showing up cleanly in web_connectivity while leaving other
test types relatively quieter.
- China (CN): no viable analysis.
Only 4 confirmed censorship incidents fall inside the
two-year window, below the 5-positive-day minimum needed to
compute AUC. This is an honest data limitation, not a model
failure — China's Great Firewall is so persistent that
most OONI days are blocked, so we lack the
blocked-vs-unblocked separation that AUC measures. The
cross-protocol classifier (
/v1/classifier/protocol/tor/CN)
is the right tool for CN questions.
Promote criteria
Both gates pass:
- ≥ 20 viable countries: we shipped 30.
- Per-country AUC range > 0.10 in at least 10 of
them: 13 countries qualify.
Honest caveats
- OONI test types have very different probe densities across
countries. web_connectivity has 22,133 rows in the
two-year window; telegram has 2,152. AUC is not
perfectly apples-to-apples — denser test types have more
statistical power.
- “Diagnostic” does not mean
“causal”. A test type can be high-AUC because
operators chose to test it heavily on incident days (e.g.
OONI volunteers re-running web_connectivity when they see
a protest brewing).
- AUC is computed in-sample (no train/test
split) because per-country sample sizes are already at the
floor. These numbers are upper bounds on out-of-sample
diagnostic power.
- The label window is ± 1 day around incidents,
which can leak signal between adjacent days and inflate
AUC for slow-burning shutdowns.
- Logistic regression uses class_weight='balanced'
because positive days are rare; an unbalanced fit would
collapse to predicting “not blocked” for
every day and report AUC 0.5.
Where the endpoints live
GET /v1/atlas/ooni-test-diagnostic/info —
global aggregates, methodology, honest caveats, promote
state.
GET /v1/atlas/ooni-test-diagnostic —
ranked countries (?limit, ?test_type, ?sort).
GET /v1/atlas/ooni-test-diagnostic/<cc> —
per-country test-type AUCs.
Sidecar at
/opt/voidly-ai/ml-deploy/ooni_test_type_diagnostic_v1.json;
build script at
scripts/build-ooni-test-type-feature-importance.py.