An observatory that asks you to trust its censorship verdicts owes you an honest answer to a prior question: how good is the data behind any one verdict? It is not uniform, and pretending otherwise would be the same kind of dishonesty this project calls out in inflated ML scores. This finding audits Voidly's own measurement coverage across every country and ships the audit as a live API so the answer travels with the data.

Not every country is measured equally

Voidly has evidence for 215 countries, but recency varies by five orders of magnitude. Most are current — 13 countries were last measured within 24 hours, 174 within the last week. But 28 countries have not been measured in over 30 days: Djibouti's most recent data point is 158 days old, built on just two measurements. A “high-risk” or “blocked” verdict for a country in that tail is a statement about the past, not necessarily the present — and a journalist citing it should know that before it goes in print.

Freshness bandCountriesMeaning
live (≤24h)13measured today
recent (≤7d)174current
stale (>30d)28weight with caution

Recency isn't the only axis — volume, stability, sources matter too

A country can be measured today and still be thinly or erratically sampled. Over the last 30 days we score each country's sampling stability as the coefficient of variation of its daily measurement counts: 54 countries are “bursty” (their signal arrives in occasional spikes rather than a steady stream) versus 24 genuinely stable. Add measurement volume and source diversity and you get a single composite data-confidence score (0–100) per country. The distribution is sobering and honest:

Confidence bandCountries (last 30d)
high (≥70)53
medium (40–69)52
low (<40)82

Iran scores 74 (high): fresh, ~3,500 measurements across three sources in 30 days — its bursty daily pattern is outweighed by sheer volume and corroboration, and its verdicts are trustworthy. Most of the 82 low-confidence countries are small or rarely-probed; their verdicts are real signals but deserve a wider error bar. A low score does not mean a country is uncensored — it means weight the verdict with caution.

The categorization blind spot

Honesty cuts the other way too. Of the 5,586 distinct domains Voidly has measured, about 69% (3,839) map to a Citizen Lab content category (NEWS, ANON, HUMR, GRP…) — so category-level analysis (“is this country blocking news sites specifically?”) is well-supported for most of the corpus, with a real ~31% blind spot for the untagged remainder. Any “by category” claim should still be read as a floor for the tagged subset, not a census.

Correction: an earlier version of this finding put that coverage at 0.5% — a measurement error that read the near-empty evidence.domain_category column instead of joining the 14,000-domain Citizen Lab category table; the true coverage is ~69%. An observatory that audits its own data quality has to audit that audit too — surfacing and fixing the gap is more useful than hiding it.

The audit is the product

All of this is now live and machine-readable, so the quality of the data travels with every query:

The same standard Voidly applies to its forecasts and classifiers — report the honest number, name the limitation, never dress a weak signal as a strong one — now applies to the measurement layer underneath them. If a country's data is thin or stale, the API will tell you so before you cite it.