The gap
A 200 response with an empty outageMessage passes every check the collector has. The key is accepted, the shape is right, there is nothing to fetch, the run records ok, and the heartbeat pings. If ESB ever changed the endpoint behind a still-valid key, or started returning an empty list for some other reason, collection would stop and nothing would say so. The webhook reports failures the collector can see, the heartbeat reports a collector that is not running, and this is neither.
The site's stale banner would not catch it either: the horizon is the last run that reached the feed (n_listed IS NOT NULL), and an empty list is a run that reached the feed.
What the data says
In 1,644 runs over five weeks the list has never been below 2 outages, and 86 runs listed fewer than 5. An empty list has not happened once. So a single empty run is already unusual, and a run of them is almost certainly not weather.
Proposed rule
Six consecutive runs with an empty list, three hours at the 30-minute interval, raise the webhook alarm once, with a banner in the style of alert.py's others: what was seen, that the key is accepted and the shape unchanged, and that the likely cause is ESB changing what the endpoint returns. Not on the first empty run: a calm night could plausibly produce one, and alerting on recoverable blips trains the reader to ignore the ones that matter (README § Alerting).
The count comes from the run table (the last six ok runs with n_listed = 0), so no new state is needed. It should not repeat every run once tripped; alert on the sixth and then, say, daily, or hold the alarm until a non-empty run clears it.
Deliberately not a heartbeat change: the collector is running and reaching the feed, so a dead-man's monitor should keep hearing from it. This is a new banner, not a new channel.
Where
esb_outages/poll.py: after the list is stored, check the run table and fire once past the threshold.
esb_outages/alert.py: a new banner and exit code, or reuse EXIT_SCHEMA_DRIFT (raw data still safe, response shape effectively changed).
tests/test_poll.py: six empty runs alert, five do not, a non-empty run resets.
- README § Alerting exit-code table, and a dated section in
notes/alerting.md.
The last remaining item from the collector list drawn up on 2026-09-05; the heartbeat (#41) and the storm cut-short (#42) are merged.
The gap
A 200 response with an empty
outageMessagepasses every check the collector has. The key is accepted, the shape is right, there is nothing to fetch, the run recordsok, and the heartbeat pings. If ESB ever changed the endpoint behind a still-valid key, or started returning an empty list for some other reason, collection would stop and nothing would say so. The webhook reports failures the collector can see, the heartbeat reports a collector that is not running, and this is neither.The site's stale banner would not catch it either: the horizon is the last run that reached the feed (
n_listed IS NOT NULL), and an empty list is a run that reached the feed.What the data says
In 1,644 runs over five weeks the list has never been below 2 outages, and 86 runs listed fewer than 5. An empty list has not happened once. So a single empty run is already unusual, and a run of them is almost certainly not weather.
Proposed rule
Six consecutive runs with an empty list, three hours at the 30-minute interval, raise the webhook alarm once, with a banner in the style of
alert.py's others: what was seen, that the key is accepted and the shape unchanged, and that the likely cause is ESB changing what the endpoint returns. Not on the first empty run: a calm night could plausibly produce one, and alerting on recoverable blips trains the reader to ignore the ones that matter (README § Alerting).The count comes from the
runtable (the last sixokruns withn_listed = 0), so no new state is needed. It should not repeat every run once tripped; alert on the sixth and then, say, daily, or hold the alarm until a non-empty run clears it.Deliberately not a heartbeat change: the collector is running and reaching the feed, so a dead-man's monitor should keep hearing from it. This is a new banner, not a new channel.
Where
esb_outages/poll.py: after the list is stored, check the run table and fire once past the threshold.esb_outages/alert.py: a new banner and exit code, or reuseEXIT_SCHEMA_DRIFT(raw data still safe, response shape effectively changed).tests/test_poll.py: six empty runs alert, five do not, a non-empty run resets.notes/alerting.md.The last remaining item from the collector list drawn up on 2026-09-05; the heartbeat (#41) and the storm cut-short (#42) are merged.