The Sentinel alert lead-time retrospective handed us an
uncomfortable number: over 90 days, 79.2% of
forecast-threshold alerts were false alarms — four out
of every five alerts that fired were not followed by a confirmed
censorship incident. A newsroom that gets burned four times stops
listening. This page is the fix, backtested and shipped, with the
before/after stated plainly.
This is the companion to the
alert
lead-time retrospective: that page diagnosed the problem, this
page fixes it. Both numbers are public and rebuild daily.
The diagnosis: it is not a tuning problem, it is a country problem
The retrospective already showed the real shape of the failure. The
79.2% aggregate hid a wide split: Egypt ran 3-for-3 (a perfect
hit rate), Uzbekistan 7-for-7, Pakistan 5-for-6 — while Iran
ran 0-for-7, and Turkmenistan, Cuba, Myanmar and others were
similarly all-noise. A single global alert threshold (a 7-day risk of
0.05) fires the same way for a country where the forecast works and a
country where it is pure background hum.
So we backtested four candidate fixes against the exact same 90-day
history and the same scoring rule the retrospective uses (an alert is
a true positive only if a confirmed censorship/mixed incident follows
within 14 days). The candidates were re-derived honestly: each one's
alert set is replayed from the daily forecast panel under that
candidate's rule, then scored.
The backtest — four candidates, two negatives
| Candidate | Alerts | False-alarm rate | True-positive rate |
| Baseline (global 0.05) | 170 | 80.6% | 17.6% |
| Per-country thresholds | 170 | 80.6% | 17.6% |
| Persistence gate (2 days) | 79 | 83.5% | 13.9% |
| Suppress chronically-false countries | 40 | 35.0% | 60.0% |
Two of the four candidates are honest negatives, and
we report them as findings rather than quietly dropping them:
- Per-country thresholds did not help. The
intuition — Egypt needs 0.4, Iran needs 0.7 — is
reasonable but wrong for this model. The forecast's probabilities
for censorship-heavy countries cluster just above 0.05 with no
separation between true-positive days and false-alarm days.
Raising a country's bar does not skim off the false alarms; it
kills the true positives along with them. The F1-optimal
threshold for almost every country turns out to be the global
floor, so the candidate collapses back to the baseline. A
genuinely better per-country threshold would need a
better-calibrated forecast first.
- The persistence gate did not help either.
Requiring the forecast to stay above threshold for two
consecutive days — meant to kill one-day spikes —
actually nudged the false-alarm rate up. Sentinel's
false alarms are not transient one-day blips; they are countries
that sit chronically near the threshold. Persistence drops a few
real alerts and barely touches the noise.
What shipped: downgrade the countries that are always wrong
The candidate that worked is the bluntest one. If a country's alerts
have been consistently, measurably wrong, that country is downgraded
from alert to watch: its 7-day risk
is still computed and still published at
/v1/sentinel/current_risk/{cc}, but no webhook alert
fires. A country is downgraded under one of three honest rules,
all measured on the 90-day backtest:
- Chronic false-positive — the country fired
at least 3 alerts and every single one was a false
alarm (a 0% true-positive rate). Wrong-every-time is unambiguous,
so the bar is low. This is the rule that catches Iran (0-for-3 in
the forecast-panel replay).
- No incident signal — the country had zero
confirmed censorship/mixed incidents in the entire 90-day window
yet still fired alerts. Every such alert is a false alarm by
construction. This is what catches stable democracies (Australia,
Germany, Japan, the Netherlands, the United States) that should
never trip a censorship alert in the first place.
- Low precision — the country fired at least
4 alerts and its precision was below 0.40. These countries
(Bangladesh 1-for-5, Kazakhstan 1-for-4, Vietnam 2-for-7) do get
some incidents right, but they are net-negative noise. A rate
needs more evidence than an absolute count, so this bar is set
higher.
37 of the 49 watched countries are downgraded under these rules.
That sounds aggressive — it is — but look at what is
kept: Uzbekistan (7-for-7), Egypt (3-for-3), Pakistan
(5-for-6), Russia (4-for-7), Nigeria (5-for-9). The countries where
Sentinel has earned trust keep alerting; the countries where it
has not are quieted to a watch signal.
Before and after
- False-alarm rate: 80.6% → 35.0% on the
90-day backtest — comfortably under the 50% promote
gate.
- True-positive rate: 17.6% → 60.0%. Note
this went up, not down. The honest worry with any
precision fix is that it craters recall; here it does not,
because the alerts that were removed were overwhelmingly noise.
The 24 true positives that survive are the ones in countries the
forecast actually predicts well.
- Median lead time: 4.2 → 3.9 days —
essentially unchanged. The early-warning margin is preserved.
The honest caveats
Reducing a false-alarm rate is the kind of result that is easy to
overclaim, so the limits are stated up front:
- The thresholds and the suppression list are picked
in-sample. They are optimised on the same 90-day window
they are scored on. The live forward false-alarm rate will be
somewhat worse than the 35.0% backtest figure — treat the
backtest number as an optimistic lower bound. The
/v1/sentinel/alert-quality endpoint rebuilds daily,
so the suppression list adapts as new alerts and incidents
accumulate.
- Suppression is a downgrade, not a deletion. A
downgraded country still has a fully computed forecast. The risk
number is public; only the push alert is withheld. If Iran's
forecast genuinely improves, Iran climbs back to alert
on the next rebuild automatically.
- Some suppressed countries are suppressed for the wrong
reason. Iran is chronically censored — the
forecast keeps crossing the threshold because Iran genuinely is a
high-risk environment. Its alerts score as false alarms partly
because no new confirmed incident landed in the matching
windows, which is as much an incident-coverage gap as a forecast
failure. Downgrading Iran is the right call for alert hygiene,
but it is not a statement that Iran is safe.
- Multi-signal confirmation was considered and not
shipped. Requiring a second independent signal — a
DBSCAN anomaly or a contagion-watchlist hit — to agree
before an alert fires is a sound idea. We did not ship it because
those signals are currently stored as point-in-time snapshots,
not historical time series, so there is no honest 90-day
backtest for them. It stays a candidate for a future fix once
that history is retained.
How to read it
GET /v1/sentinel/alert-quality returns the full
backtest: every candidate with its false-alarm and true-positive
rate, the winning candidate, the before/after summary, the
suppressed-country list with the rule that caught each country, and
the per-country alert thresholds.
GET /v1/sentinel/alert-quality/{cc} returns one
country's status — whether it is on alert or
watch, and why. It is paired with
/v1/sentinel/alert-lead-time, which is the
accountability number this fix is measured against. Both rebuild
daily; both report the uncomfortable parts in the open.