The bug

Voidly's ingest pipeline promotes critical IODA ASN-level outage alerts into the incidents table with incident_type='disruption' and severity='critical'. The forecast feature builder loaded ALL incidents and labeled country-day pairs positive whenever any such incident fell in the next 7 days. The unintended consequence: IODA outages — which include fiber cuts, BGP changes, DDoS attacks, weather, and routine maintenance — were being treated as confirmed censorship.

April 2026 had 1,074 incidents across 167 countries. 1,011 were IODA disruption. Only 45 were CensoredPlanet pure-censorship and 18 were mixed. So 94% of April incidents were noise from a label perspective.

The downstream damage

The fix

scripts/build-forecast-features.py patched to exclude incident_type = 'disruption' from the load_incidents query. Censorship + mixed incidents kept; IODA disruption excluded from the forecast target only. The disruption incidents themselves are NOT deleted from the table — they remain visible on country pages with their honest type label. Only the forecast training treats them differently.

Result: monthly positive rate before vs after

MonthBeforeAfter
2026-0130.6%26.4%
2026-0227.0%27.0%
2026-0360.7%16.3%
2026-0478.9%20.6%
2026-0536.3%12.2%

Model trained on sane labels — promoted to production

Retrained the forecast model on the sane labels. Dual-gate result:

Feature importance shifted: GDELT unrest dropped from 25% to 11% (over-rewarded by disruption labels), weight redistributed to recent-shutdown, rolling block-rates, and seasonal markers — a sane attribution for a censorship forecaster.

Regression test added

New scripts/test-forecast-labels-sane.py runs after feature build in the weekly retrain. Fails the pipeline if any of the last 12 months has positive rate above 40%. The 79% explosion would have been caught instantly. Wired into weekly-retrain.sh as stage test-forecast-labels-sane.

Honest scope

This fixes the FORECAST training labels. The incidents table itself still contains the IODA disruption entries; we keep them because they are real network observations and journalists care about them. We do NOT count them as “censorship incidents” in the forecast or in canonical incident-count headlines going forward. The Atlas Score v1 and v2 weights still consider disruption signal, but with much lower weight than confirmed censorship.