voidly

Benchmarks / The evidence landscape

Compare the
evidence.

Choose the source that answers
your question.

Explore the sources
QuestionTime windowUnit

A comparison needs all three.

01 / Start with the question

Different instruments.
Different answers.

Measurement volume, narrative depth and model skill are distinct. This is a source guide, not a shared accuracy leaderboard.

Provider documentation reviewed . Source populations and terms can change.

01VoidlyEvidence + incident recordsWhat does the evidence connect to?
Corpus
An aggregation layer. Measurement rows, linked evidence and incident records are different units; do not sum overlapping upstream corpora.
Coverage
Read the index’s country rows and observation window. The bundled all-time sample total and scored-country list have different coverage.
Freshness
Use dataAsOf and per-country source dates. generated_at_utc and dateModified on the index are response times. Published schedules are not proof of a completed ingest.
API access
Public JSON/CSV and incident endpoints. Authentication and limits depend on the endpoint.
Citation
Stable incident pages and BibTeX/RIS reports. Total incident records include suspected and disruption records; a stable ID is not a censorship determination.
Classification
Versioned classifier evaluation and separate forecast targets. Inspect the model record below and its label-composition caveat.
Agent access
Published @voidly/mcp-server and @voidly/agent-sdk. Inspect the package version and supported host instructions; no tool-count ranking.
License
Voidly-original material: CC BY 4.0. Upstream measurements and test lists retain their own terms.
02OONINetwork measurementsWhat happened in this network test?
Corpus
Probe test results, with test-specific methodology. The underlying corpus is not an independent addition to a downstream subset.
Coverage
Depends on volunteer tests, networks, countries and the period queried. Country presence is not representative national coverage.
Freshness
Inspect individual measurement timestamps and the selected Explorer/API time range.
API access
Public OONI API and Explorer; published measurement data can be inspected directly.
Citation
Measurement URLs and research reports. Cite the test and time range supporting the claim.
Classification
Test-specific anomaly and blocking analysis. This is not the same task as Voidly’s country-day classifier.
Agent access
API access is documented. This page makes no unsupported claim that no MCP integration exists.
License
Published OONI data: CC BY-NC-SA 4.0; check the data license and any source-specific notices.
03Freedom HouseCountry reports + human scoringWhat is the broader freedom context?
Corpus
Freedom on the Net country reports and scores; not a corpus of network-test rows.
Coverage
Read the country selection and reporting period for the named annual edition.
Freshness
Annual reporting periods, not a streaming observation feed.
API access
Use the report and data-download options for the specific edition. A downloadable report does not imply a general-purpose API.
Citation
Cite the report edition, country chapter and reporting period.
Classification
Human-scored internet-freedom indicators on a 0–100 scale. Not directly comparable with classifier F1.
Agent access
Use source documents and any documented data interfaces; no agent-tool ranking is inferred.
License
Check the report and dataset’s published reuse terms. Free to read does not establish unrestricted redistribution.
04Cloudflare RadarTraffic, attacks + outagesHow is observed internet traffic changing?
Corpus
Aggregated observations from Cloudflare’s network and public DNS resolver. Different visibility from on-device tests.
Coverage
Depends on the metric, location, time range and traffic visible to Cloudflare.
Freshness
Read each endpoint’s aggregation interval, date range and confidence. “Real time” is not one universal latency.
API access
Free Radar API; follow its authentication and request-limit documentation.
Citation
Share charts and use published outage material with the selected time range. The old “no curated incidents” label is not a safe general claim.
Classification
Traffic and attack indicators, not a matched evaluation of the Voidly classifier target.
Agent access
Cloudflare publishes a Radar MCP server. The old table’s “no public MCP server” entry is superseded.
License
Radar API data: CC BY-NC 4.0, as stated in its current overview.
05M-LabPerformance test dataHow did the connection perform?
Corpus
Raw output from measurement tests, with a subset parsed into BigQuery. Test rows are not curated censorship incidents.
Coverage
Choose the test, date range and relevant network population. Do not infer national coverage from server presence.
Freshness
The data overview describes typically at least a 24-hour collection-to-publication delay. The old “real-time BigQuery” claim is superseded.
API access
Google Cloud Storage archives and BigQuery, depending on the test. Check query costs and applicable cloud quotas separately.
Citation
Cite the test dataset, date range and source URL or table.
Classification
Network performance measurement. This comparison does not establish censorship-classifier accuracy.
Agent access
Published data interfaces and SQL; no unsupported claim that no MCP integration exists.
License
Data collected by M-Lab tests: CC0. Website material has separate terms.
06Citizen LabResearch + test listsWhich targets and context need investigation?
Corpus
Investigative reports and categorized URL test lists. A listed URL is a test target, not a measurement or proof of a block.
Coverage
Investigation scope and country/global test lists differ. The global list is not a uniform sample of every country.
Freshness
Use the report date or repository revision. Reports and list changes are not a continuous measurement stream.
API access
Public reports and GitHub lists, including CSV/JSON. No continuous incident API is claimed here.
Citation
Cite the investigation or the test-list repository/revision.
Classification
Investigative methods and target categorization, not a matched country-day classifier.
Agent access
Read the relevant publications and repository. This page does not assert an exhaustive absence of integrations.
License
Test-list data: CC BY-NC-SA 4.0. Research reports may have different artifact-specific terms.

02 / Model evidence

A score needs
a test.

Country-day classification, future event risk and network performance answer different questions. Their metrics do not share a ranking.

Classifier / v3.3Source loaded

GradientBoostingClassifier

LOCO mean F10.711127 evaluated countries
Well-sampled mean F10.63061 countries with n ≥ 30

794 of 1,116 positive training labels (71.1%) are country-days whose only incident is an IODA `disruption` row — network outages, not confirmed censorship. The same class was excluded from forecast labels in 2026-05 as ~94% noise. So this model is substantially trained to detect DISRUPTION, and its F1 should be read as such. Corpus is frozen at 2026-05-21; relabelling is a pending decision, not an oversight.

Model trained
2026-05-21T03:01:46.793987+00:00
Record generated
Not supplied
Request checked

Cross-country F1 is not a per-prediction probability or a future-time accuracy estimate. Loaded introspection does not establish which model every prediction route serves.

Evaluation population, distribution & source caveats
Source headline
LOCO mean F1 across all countries
Training rows
4,237
Positive training rows
1,116
Training countries
131
LOCO median F1 · distribution caveat
0.870
Countries with perfect F1
46
Stratified F1 · different split
0.729
Stratified AUC · different metric
0.899

The headline LOCO MEDIAN F1 (0.870) is dominated by many small-sample countries scoring a perfect 1.0 on a handful of days; the MEAN F1 (0.711) is the honest single number. Censorship-heavy, high-volume countries score materially lower (CN ~0.29 on n=95, BY ~0.21 on n=84, AZ ~0.11 on n=65). Cite the mean — or the specific per-country number — for hard countries, not the median. The model is CLEAN (no label leakage); this is a distribution caveat, not an accuracy retraction.

HONESTY: forward-temporal holdout (train past, test future) gives AUC 0.669 / F1 0.474 vs the random-split AUC 0.895 / F1 0.725. v3.3 generalizes across COUNTRIES (LOCO F1 0.87) but DEGRADES across TIME (delta AUC -0.226) — which is why it is retrained weekly. Do not read the random-split number as forward-deployment accuracy. See /atlas/findings/classifier-v3.3-temporal-generalization-2026-05.

confidence is

stratified cross-validation F1, NOT a calibrated per-prediction probability

correct use

Cross-country censorship-risk classification. Per-country accuracy varies widely — 16 MENA / former-Soviet countries (OM, UZ, TN, LY, YE, JO, MA, …) regress 5-29pp on sparse neighbor-pair overlap. See evaluation.honest_evaluation + evaluation.per_country for the full distribution.

headline metric

LOCO median F1 0.87 (cross-country); honest MEAN F1 0.71; well-sampled (n>=30 countries) mean ~0.63

model status

CLEAN — no label leakage (ML_LEAKAGE_AUDIT.md); this caveat is about the score distribution, not leakage.

training label composition

794 of 1,116 positive training labels (71.1%) are country-days whose only incident is an IODA `disruption` row — network outages, not confirmed censorship. The same class was excluded from forecast labels in 2026-05 as ~94% noise. So this model is substantially trained to detect DISRUPTION, and its F1 should be read as such. Corpus is frozen at 2026-05-21; relabelling is a pending decision, not an oversight.

Different historical random-split runs report different values. Keep the source version and split with each number; do not merge them into one evaluation.

Temporal-generalization study ↗

Where Voidly fits

Keep the
source trail.

Voidly combines evidence from OONI, CensoredPlanet, IODA, Citizen Lab test lists and its own probes into navigable records. An incident ID connects a record to its evidence and exports.

OONI supplies test-level observations. Radar offers its network’s traffic view. Freedom House brings country research and human coding. M-Lab focuses on performance. Citizen Lab supplies investigations and target context. Voidly is complementary; it does not replace their source work.

The classifier, agent interfaces and seven-day forecast are separate capabilities. A forecast response is not proof of a coming shutdown, and a shared URL is not independent corroboration.

03 / Follow the record

Open the evidence.

Read-only API examples. These requests inspect records; they do not reproduce every provider comparison or train a model.

The example IR-2026-0142 is an IODA disruption record in the source checked for this page. Its identifier alone does not establish intentional censorship. Inspect its type, dates and linked evidence.

cURL examples
# 1. Country snapshot: the response is an object with a countries array
curl -fsS 'https://api.voidly.ai/data/censorship-index.json' | jq '.countries[0]'

# 2. Recent Iran incident records (including disruption/suspected records)
curl -fsS 'https://api.voidly.ai/data/incidents?limit=10&country=IR' | jq '.incidents[].hashId'

# 3. Inspect this example record before drawing a conclusion
curl -fsS 'https://api.voidly.ai/data/incidents/IR-2026-0142' | jq .

# 4. Inspect linked evidence and its source
curl -fsS 'https://api.voidly.ai/data/incidents/IR-2026-0142/evidence' | jq '.evidence[0]'

# 5. Seven-day route: retain model identity and interpretation notes
curl -fsS 'https://api.voidly.ai/v1/forecast/IR/7day' | jq '{model_version, summary, honest_forecast, honest_caveat, forecast_trajectory_note}'
JavaScript example & host setup
JavaScript example
// Modern Node.js provides fetch; no extra package is required.
const response = await fetch(
  'https://api.voidly.ai/data/incidents/IR-2026-0142',
  { signal: AbortSignal.timeout(8000) }
);
if (!response.ok) throw new Error(`Incident request failed: ${response.status}`);
const incident = await response.json();
console.log(incident.title, incident.severity, incident.hashId);
console.log(incident.incidentType, incident.sources, incident.status);

// MCP host setup: /integrations
// npx @voidly/mcp-server

The native fetch example handles HTTP errors and has an eight-second timeout. Host-specific MCP configuration is in Integrations.

04 / Leave a reference

Cite the right unit.

Name the dataset or incident, observation period, model version where relevant, and access date. Keep upstream terms with the material.

General BibTeX reference
general BibTeX
@misc{voidly2026,
  author       = {{Voidly}},
  title        = {{Voidly: Censorship intelligence + agent
                   infrastructure for the post-AI internet}},
  year         = {2026},
  howpublished = {\url{https://voidly.ai}},
  note         = {Voidly original material: CC BY 4.0;
                  upstream data retains its source license.
                  Record the version and time range used.}
}

This preserves the original general-reference title and year; it omits mutable corpus counts from an evergreen reference. For publication dates and other formats, use the citation hub.

Single-incident BibTeX example
incident BibTeX
@misc{voidly_IR_2026_0142,
  author = {{Voidly}},
  title  = {{Incident record IR-2026-0142}},
  year   = {2026},
  url    = {https://voidly.ai/incident/IR-2026-0142},
  note   = {Example record; inspect incident type and source evidence.
            Voidly original material: CC BY 4.0; upstream terms apply.}
}

The example is labeled as an incident record rather than asserting its cause. Original record · Original legacy destination.

RIS / API export

Public incident reports export BibTeX and RIS. The export’s source conventions may use observation-year and API URLs; save the exact file with your work.

citation export commands
curl 'https://api.voidly.ai/data/incidents/IR-2026-0142/report?format=bibtex'
curl 'https://api.voidly.ai/data/incidents/IR-2026-0142/report?format=ris'

Voidly-original material is CC BY 4.0. All third-party material retains its own terms. The published FACTS.md is a reference document; its dated figures do not become current through this page.

Reference history

What changed in
the comparison.

The old numbers remain available for review below. They are not current measurements, a harmonized benchmark, or a basis for provider rankings.

Previously published comparison · original values & corrections

Captured from the original page source on 2026-09-06. That is a capture date, not the date the measurements were made. Claims inside this archive may be stale or incorrect. Use the source guide and corrections instead of copying them as current facts.

Coverage & volume

The original table mixes all-time samples, historical archives, live windows, lifetime evidence and provider-wide totals. It does not establish deduplicated combined volume or comparable national coverage.

Freshness

A rebuild every day does not refresh hardcoded source figures. M-Lab documents a typical delay of at least 24 hours; the old real-time label is superseded. The old Voidly ingest/probe intervals are published schedules, not current run evidence.

Incidents & sources

The total incident count is not a count of citable censorship events. Some records describe disruption or suspected signals. The old ~100× OONI depth ratio and “unmatched” Radar assertion have no matched population/evaluation on this page.

Access & licensing

The old OONI CC0/CC BY-SA, M-Lab CC BY-SA and Citizen Lab test-list CC BY 2.5 entries are superseded by their linked data licenses above. Radar explicitly publishes API data under CC BY-NC 4.0. Voidly’s license does not override upstream terms.

Agent integration

The original page conflicts between 84 and 125 tools. Neither number is a current capability benchmark. Cloudflare publishes a Radar MCP server, so the old absence claim and REST-only ranking are superseded.

Bar lengths

All old bars were hand-assigned relative indicators, explicitly not precise scores. Their numeric widths below are archival design values, not measurement results or a defensible ranking.

Original overview · eight dimensions

Previously published: Total measurements (corpus)
Voidly
38.8M all-time + 1.6M historical

Aggregated from OONI/CensoredPlanet/IODA + Voidly probe network. Live = last 30 days.

OONI
3.5B+ (OONI)

Underlying OONI raw measurement corpus (Voidly aggregates from this).

Freedom House

Not a measurement dataset — annual qualitative + scored report.

Cloudflare Radar
see docs

Massive, real-time. Public counters but no per-row export of the underlying dataset.

M-Lab
see docs

Open via BigQuery — billions of NDT/MSAK/etc. rows.

Citizen Lab

Publishes test lists (the global list is 1,725 URLs; Voidly has ingested 5,154 distinct domains across the country lists) + reports, not a continuous measurement feed.

Previously published: Countries covered
Voidly
130 live, 168 in evidence DB

Live = countries with current OONI / probe samples; evidence DB = lifetime corpus.

OONI
195+

Global probe coverage; gaps where probes are not deployed.

Freedom House
~70

FOTN 2024 covered 72 countries with full reports.

Cloudflare Radar
200+

Effectively every country/region with Cloudflare traffic.

M-Lab
200+

Worldwide via the M-Lab measurement platform.

Citizen Lab
see reports

Coverage varies per investigation.

Previously published: Freshness (latest data)
Voidly
6h aggregate, 5min probes

OONI/IODA/CensoredPlanet ingest runs every 6h; Voidly probe network reports every 5min.

OONI
minutes

Probe submissions appear on Explorer near-real-time.

Freedom House
annual

Freedom on the Net publishes once per year.

Cloudflare Radar
real-time

Live counters; some indicators are 5-min rolling.

M-Lab
real-time

BigQuery tables update continuously.

Citizen Lab
periodic

Investigative reports — episodic, not continuous.

Previously published: Free public API
Voidly
Yes

CC BY 4.0, no auth on public endpoints, rate-limited.

OONI
Yes

Public, well-documented JSON API + S3 raw.

Freedom House
Partial

PDFs + CSV download free; richer datasets are gated.

Cloudflare Radar
Yes (limited)

Radar API requires Cloudflare account; rate-limited by tier.

M-Lab
Yes

BigQuery public datasets — Google Cloud egress costs apply.

Citizen Lab

Reports are public; no continuous API feed.

Previously published: Citable incident IDs
Voidly
Yes (7,500+)

Stable human-readable IDs e.g. IR-2026-0142. BibTeX/RIS export.

OONI
measurement URLs

Per-measurement permalinks; no curated incident layer.

Freedom House
cite report

Cite the country chapter / annual report.

Cloudflare Radar
no

Charts and snapshots; no citable incident records.

M-Lab
measurement IDs

Per-test rows in BigQuery; no incident curation layer.

Citizen Lab
cite report

Cite individual investigations.

Previously published: ML classification / scoring
Voidly
0.71 F1 (v3.3 LOCO mean)

GradientBoosting v3.3, LOCO MEAN across 127 countries — 0.63 across the 61 with n≥30. The 0.87 median often quoted is inflated by 46 small-sample countries scoring a perfect 1.0 on a handful of days; /v1/classifier/info names the mean as its headline metric. v2 reported 99.8% F1 but was retired 2026-05-21 (country-tier leakage).

OONI
rule-based

Deterministic anomaly detection per nettest.

Freedom House
0–100 score

Human-coded across 21 indicators.

Cloudflare Radar

Traffic + threat indicators, not censorship-classified.

M-Lab

Performance-focused; no censorship label.

Citizen Lab

Manual investigative methodology.

Previously published: MCP / AI agent integration
Voidly
84 tools

@voidly/mcp-server + @voidly/agent-sdk on npm.

OONI
no

API only — no MCP server published.

Freedom House
no

PDF-first publication.

Cloudflare Radar
no

No public MCP server.

M-Lab
no

No public MCP server.

Citizen Lab
no

No public MCP server.

Previously published: License
Voidly
CC BY 4.0

All published Voidly datasets.

OONI
CC0 / CC BY-SA

Raw measurements CC0; some derived works CC BY-SA.

Freedom House
mixed

Reports free to read; redistribution requires permission.

Cloudflare Radar
see ToS

Cloudflare ToS governs Radar API output.

M-Lab
CC BY-SA 4.0

Open data, share-alike attribution.

Citizen Lab
CC BY 2.5 CA

Most reports/test lists; check per artifact.

Original depth slices · editorial indicators

Previously published: Country coverage
Archived values / countries. Indicator widths are arbitrary, not scores.
ProviderOld valueOld width (%)
Cloudflare Radar200+100
M-Lab200+100
OONI195+98
Voidly130 live / 168 lifetime84
Freedom House~7035
Previously published: Data freshness
Archived values / latency. Indicator widths are arbitrary, not scores.
ProviderOld valueOld width (%)
Cloudflare Radarreal-time100
M-Labreal-time100
OONIminutes95
Voidly5min probes / 6h aggregate75
Citizen Labperiodic reports25
Freedom Houseannual5
Previously published: Tracked incident records
Archived values / curated incident records. Indicator widths are arbitrary, not scores.
ProviderOld valueOld width (%)
Voidly7,500+ (stable IDs)100
OONImeasurement URLs only60
Citizen Labinvestigation reports40
Freedom Houseannual chapter cite25
M-Labmeasurement rows50
Cloudflare Radarno curated incidents5
Previously published: Accessibility (free + permissive license)
Archived values / openness. Indicator widths are arbitrary, not scores.
ProviderOld valueOld width (%)
OONICC0 / CC BY-SA — free API100
VoidlyCC BY 4.0 — free API95
M-LabCC BY-SA — BigQuery (egress fees)80
Citizen LabCC BY 2.5 CA — reports + test lists70
Cloudflare Radarfree API w/ account, ToS-bound55
Freedom HousePDF free, datasets gated40
Previously published: AI / MCP integration
Archived values / agent integration depth. Indicator widths are arbitrary, not scores.
ProviderOld valueOld width (%)
Voidly125 MCP tools + agent SDK100
OONIREST API only25
Cloudflare RadarREST API only25
M-LabBigQuery / SQL only20
Freedom HousePDF-first5
Citizen LabPDF / GitHub10

The earlier “where we’re unique” paragraph claimed 5+ source networks and a 30+ node probe fleet from a configuration constant, alongside 125 MCP tools, v3.3 LOCO F1 and XGBoost/isotonic seven-day forecasting. Those historical fleet/tool figures are not current operational evidence. The original “no other” agent-consumption claim was not an exhaustive ecosystem evaluation.

Retired classifier v2’s 99.8% F1 was attributed to country-tier leakage and retired on 2026-05-21. It is not a current score. The original v3.3 mean, well-sampled subset, small-sample median caveat and retirement date remain in the archived classification row; the current source card adds the training-label and forward-time caveats.

The old page called its bars rough relative indicators, not precise scores, and said to consult source documentation where a number was not disclosed. That limitation remains. It also claimed every Voidly number could be reproduced with its public API examples; those examples do not reconstruct historical classifier evaluation or every comparison.

Original comparison reference ↗