Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
18 commits
Select commit Hold shift + click to select a range
13102dc
Deliver alerts best effort, and never print their URLs
claude Sep 24, 2026
2c12d33
Alert when a run fails partway, not only when the probe does
claude Sep 24, 2026
9d554db
Refuse a negative or non-numeric poll delay
claude Sep 24, 2026
aa5ba05
Do not raise a storage alarm when an overlapping run removed the probe
claude Sep 24, 2026
80de7d7
Read only a held lock as held, and wait out a brief holder
claude Sep 24, 2026
5674cd1
Test the alert paths nothing held
claude Sep 24, 2026
2bf8d73
Name the 30-minute cadence and the directory in use in the storage fix
claude Sep 24, 2026
b3d16c7
Cap the error text an alert can carry
claude Sep 24, 2026
effffc5
Leave room under the backstop for the last fetch, the webhook and the…
claude Sep 24, 2026
1f5e8c8
Let rebuild and compact work beside a malformed database
claude Sep 24, 2026
b5099ca
Announce every backup failure, not only fetch, merge and push
claude Sep 24, 2026
cf60cc9
Keep the backup's git errors in variables, not fixed files in /tmp
claude Sep 24, 2026
61a2061
Take the poll lock while the backup commits and merges
claude Sep 24, 2026
0c96288
Time out a stalled backup, and say so when it does
claude Sep 24, 2026
90474cd
Pass the wrapper's secrets to sudo in the environment, not as arguments
claude Sep 24, 2026
4b2bb30
Gate the interpreter the service actually runs
claude Sep 24, 2026
5993d10
Stop the install on an empty host key scan, and pin the key it seeds
claude Sep 24, 2026
0fbc9cb
Document the backup setup where it is pointed to
claude Sep 24, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -117,8 +117,8 @@ Every one of these has already cost someone an hour:
| Part-observed days keep their colour and say so in the tooltip | `notes/grading.md` § Short days say so |
| 2.5M customer denominator, and which DAPR figures are comparable | `notes/grading.md` § The customer denominator |
| Each run logs two lines to `runs-*.jsonl`: the list before any detail is fetched, and an `"event": "end"` line (`Store.finish_run`) carrying what a rebuild cannot derive - status, exit code, finish time, skipped count, errors. Older runs have no end line and replay from the start line alone | `notes/storms.md` § The run's end is logged after all (2026-09-24) |
| The collector pings `ESB_HEARTBEAT_URL` after every run that reached the feed, drifted or partial ones included, and never after a rejected key, an unreachable feed or a skipped trigger; a dead-man's monitor alerts on silence (period 30 min, grace 90). `test-alert` proves both channels | `notes/alerting.md` § The heartbeat (2026-09-06) |
| A storm can list more than a run can fetch: `ids_needing_detail` ranks by what a purge would take (listed Restored and still needing a fetch, then live never-fetched, then re-checks), every detail is committed as it lands, a run stops itself at `RUN_BUDGET_S` (24 min) and records `cut_short` (exit 0, heartbeat sent, no webhook), SIGTERM from the service unit's backstop ends the loop the same way, and the pause is 500 ms. In list order a cut-short run re-fetched the same head every run and never reached the tail | `notes/storms.md` (2026-09-06) |
| The collector pings `ESB_HEARTBEAT_URL` after every run that reached the feed, drifted or partial ones included, and never after a rejected key, an unreachable feed, a storage failure, a crash or a skipped trigger; a dead-man's monitor alerts on silence (period 30 min, grace 90). `test-alert` proves both channels | `notes/alerting.md` § The heartbeat (2026-09-06); § A run that fails partway (2026-09-24) |
| A storm can list more than a run can fetch: `ids_needing_detail` ranks by what a purge would take (listed Restored and still needing a fetch, then live never-fetched, then re-checks), every detail is committed as it lands, a run stops itself at `RUN_BUDGET_S` (22 min, four short of the backstop) and records `cut_short` (exit 0, heartbeat sent, no webhook), SIGTERM from the service unit's backstop ends the loop the same way, and the pause is 500 ms. In list order a cut-short run re-fetched the same head every run and never reached the tail | `notes/storms.md` (2026-09-06) |
| The Pi pushes every six hours and `esb-data` dispatches the site build on each push; the two crons are a fallback only, because scheduled runs here have landed 4-10h behind their cron time. The stale banner trips at 10h - above the widest legitimate age (~7h), below a missed push (13h+) | `notes/publish-cadence.md`; `STALE_AFTER` in `esb_site/render.py` |
| The banner states the data's *age* ("Updated 17 hours ago"), not its timestamp, and names no cause: from the browser a stalled build and a stalled collector look identical. A healthy overnight gap is a big number, so the warning, not the wording, carries it | `freshness()` in statusui's `ui.js`; `notes/publish-cadence.md` § The banner blamed the wrong half |
| The exact horizon left the footer (owner call, 2026-08-26) and then the county page's header (2026-08-28); it survives only as the age chip's hover title on the index. A county page's currency signal is the month table's `to 27 Aug` caveat | `notes/design-alignment.md` § The county page became an archive |
Expand Down
5 changes: 3 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -32,7 +32,7 @@ change of outage type forces an immediate detail fetch however long that outage
has been dormant. Only a quiet outage's descriptive fields are ever delayed.

A storm can list more outages than one run can fetch. A run stops itself at
24 minutes, about 2,800 details, records itself as cut short, and the next run
22 minutes, about 2,550 details, records itself as cut short, and the next run
fetches first whatever a purge would take: outages listed as restored that
still need their detail, then live ones never seen, then re-checks. Every detail
is committed as it lands, so even a run killed outright leaves the database
Expand Down Expand Up @@ -128,11 +128,12 @@ banner it prints:
| Exit | Meaning |
| --- | --- |
| 0 | Success — silent |
| 1 | The run crashed on an error it has no handling for; the traceback is in the journal |
| 2 | **API key rejected (HTTP 401)** — collection has stopped |
| 3 | ESB API unreachable after retries |
| 4 | API response shape changed (raw data still safe) |
| 5 | A broad failure of detail fetches |
| 6 | Data directory not writable |
| 6 | Data directory not writable, before the run or during it (a full disk) |

Deliberately *not* alerts: a per-outage 404 (the outage was purged between the
list call and its detail call — routine), and one or two isolated fetch failures.
Expand Down
30 changes: 20 additions & 10 deletions esb_outages/__main__.py
Original file line number Diff line number Diff line change
Expand Up @@ -3,13 +3,14 @@
from __future__ import annotations

import argparse
import contextlib
import os
import sys
from pathlib import Path

from . import __version__, alert
from .client import EsbClient
from .poll import DEFAULT_DELAY_MS, poll_lock, run_check, run_poll
from .poll import DEFAULT_DELAY_MS, milliseconds, poll_lock, run_check, run_poll
from .store import Store

DEFAULT_DATA_DIR = os.environ.get("ESB_DATA_DIR", "/data")
Expand Down Expand Up @@ -104,17 +105,23 @@ def cmd_test_alert(args) -> int:
return alert.EXIT_OK


def _held_by_poll() -> int:
def _lock_held() -> int:
# Both delete files a poll writes to: esb.db and its journal, or the log.
print("a poll run holds the lock; try again when it has finished", file=sys.stderr)
print(
"the lock is held (by a poll, the backup, or esb rebuild or compact);"
" try again when it has finished",
file=sys.stderr,
)
return 1


def cmd_rebuild(args) -> int:
with poll_lock(Path(args.data_dir)) as acquired:
if not acquired:
return _held_by_poll()
with Store(args.data_dir) as store:
return _lock_held()
# Not opened first: a malformed esb.db fails the moment it is opened,
# and rebuild deletes it unread.
with contextlib.closing(Store(args.data_dir)) as store:
result = store.rebuild(verbose=True)
if result["runs"] == 0 and result["observations"] == 0:
print("nothing to replay: no raw logs found", file=sys.stderr)
Expand All @@ -124,9 +131,9 @@ def cmd_rebuild(args) -> int:
def cmd_compact(args) -> int:
with poll_lock(Path(args.data_dir)) as acquired:
if not acquired:
return _held_by_poll()
with Store(args.data_dir) as store:
done = store.compact()
return _lock_held()
# Only raw/, so not opened: a malformed esb.db must not stop it.
done = Store(args.data_dir).compact()
print(f"compacted {len(done)} file(s): {', '.join(done) or 'none'}")
return alert.EXIT_OK

Expand All @@ -143,8 +150,11 @@ def main(argv=None) -> int:

p_poll = sub.add_parser("poll", help="run one collection pass (the scheduled command)")
p_poll.add_argument(
"--delay-ms", type=int, default=None,
help=f"pause between detail requests (env: ESB_POLL_DELAY_MS, default {DEFAULT_DELAY_MS})",
"--delay-ms", type=milliseconds, default=None,
help=(
"pause between detail requests, 0 or more "
f"(env: ESB_POLL_DELAY_MS, default {DEFAULT_DELAY_MS})"
),
)
sub.add_parser("check", help="verify the API key and connectivity; writes nothing")
sub.add_parser(
Expand Down
62 changes: 53 additions & 9 deletions esb_outages/alert.py
Original file line number Diff line number Diff line change
Expand Up @@ -12,9 +12,11 @@
import json
import os
import sys
import urllib.error
import urllib.request

EXIT_OK = 0
EXIT_CRASH = 1
EXIT_AUTH = 2
EXIT_UNREACHABLE = 3
EXIT_SCHEMA_DRIFT = 4
Expand All @@ -23,6 +25,7 @@

EXIT_MEANINGS = {
EXIT_OK: "success",
EXIT_CRASH: "collector crashed",
EXIT_AUTH: "API subscription key rejected",
EXIT_UNREACHABLE: "ESB API unreachable",
EXIT_SCHEMA_DRIFT: "API response shape changed",
Expand All @@ -33,6 +36,12 @@

BANNER_WIDTH = 78

DELIVERY_TIMEOUT_S = 10

# Discord rejects a message over 2,000 characters outright; ntfy allows 4,096.
MAX_ALERT_CHARS = 1900
TRUNCATED = "\n[truncated; the full text is in the journal]"


def banner(title: str, lines: list[str]) -> str:
# Fixed width: these end up in an email, and a long raw error message would
Expand Down Expand Up @@ -71,7 +80,7 @@ def unreachable_banner(detail: str) -> str:
"ESB POLLER: API UNREACHABLE",
[
"The outage list endpoint could not be reached after retries.",
"If this clears on the next hourly run, no action is needed - a",
"If this clears on the next run, no action is needed - a",
"single miss is covered by the ~4h retention window. Repeated",
"failures mean data is being lost.",
"",
Expand Down Expand Up @@ -100,13 +109,31 @@ def storage_banner(data_dir, problem: str) -> str:
[
f"{problem}",
"",
"Nothing was collected. Usual causes are a full disk, or the",
"Collection has stopped. Usual causes are a full disk, or the",
"directory not being owned by the user the collector runs as.",
"",
"Check:",
f" df -h {data_dir}",
f" ls -ld {data_dir}",
" sudo chown -R esb:esb /var/lib/esb-outages",
f" sudo chown -R esb:esb {data_dir}",
],
)


def crash_banner(exc: BaseException) -> str:
return banner(
"ESB POLLER: RUN CRASHED",
[
"The collector hit an error it has no handling for and stopped",
"partway through the run. The traceback is in the journal:",
" journalctl -u esb-outages.service -n 50",
"",
"If the database is at fault, this re-derives it from the raw",
"logs: sudo esb rebuild",
"If the rebuild fails the same way, or the next run crashes",
"again, the code needs a fix; rebuild again once it has one.",
"",
f"Raw error: {type(exc).__name__}: {exc}",
],
)

Expand All @@ -125,29 +152,46 @@ def partial_banner(failed: int, attempted: int, errors: list[str]) -> str:
)


def _deliver(request, what: str) -> bool:
def _deliver(what: str, url: str, data: bytes | None = None, headers=None) -> bool:
"""Best effort, in one place: a failure to report must never mask the
problem being reported or change the exit code."""
try:
urllib.request.urlopen(request, timeout=10).close()
# Built in here: a URL missing its scheme raises from the constructor.
request = urllib.request.Request(url, data=data, headers=headers or {})
urllib.request.urlopen(request, timeout=DELIVERY_TIMEOUT_S).close()
return True
except Exception as exc:
print(f"warning: {what} failed: {exc}", file=sys.stderr)
print(f"warning: {what} failed: {_describe(exc)}", file=sys.stderr)
return False


def _describe(exc: Exception) -> str:
# Never str(exc): the URL is the secret (an ntfy topic, a ping id), and
# urllib and http.client quote it, or its path, in their messages.
if isinstance(exc, urllib.error.HTTPError):
return f"HTTP {exc.code}"
reason = exc.reason if isinstance(exc, urllib.error.URLError) else exc
if isinstance(reason, OSError):
return f"{type(reason).__name__}: {reason.strerror or 'no detail'}"
if isinstance(reason, ValueError):
return f"{type(reason).__name__} (check the URL's form)"
# A URLError's reason can be a bare string, which may quote the URL.
return type(reason if isinstance(reason, Exception) else exc).__name__


def notify(message: str) -> bool:
"""Push to ESB_ALERT_WEBHOOK. Returns whether it was delivered."""
url = os.environ.get("ESB_ALERT_WEBHOOK")
if not url:
return False
if len(message) > MAX_ALERT_CHARS:
message = message[: MAX_ALERT_CHARS - len(TRUNCATED)] + TRUNCATED
if "ntfy" in url:
data, headers = message.encode("utf-8"), {"Title": "ESB poller failure"}
else:
data = json.dumps({"content": message, "text": message}).encode("utf-8")
headers = {"Content-Type": "application/json"}
req = urllib.request.Request(url, data=data, headers=headers, method="POST")
return _deliver(req, "alert webhook")
return _deliver("alert webhook", url, data, headers)


def heartbeat() -> bool:
Expand All @@ -160,7 +204,7 @@ def heartbeat() -> bool:
url = os.environ.get("ESB_HEARTBEAT_URL")
if not url:
return False
return _deliver(url, "heartbeat ping")
return _deliver("heartbeat ping", url)


def fail(message: str, code: int) -> int:
Expand Down
5 changes: 5 additions & 0 deletions esb_outages/client.py
Original file line number Diff line number Diff line change
Expand Up @@ -33,6 +33,7 @@

DEFAULT_TIMEOUT = 15.0
DEFAULT_RETRIES = 3
ERROR_BODY_CHARS = 200


class EsbError(Exception):
Expand Down Expand Up @@ -115,6 +116,10 @@ def _request(self, url: str) -> str:
body = _decode(exc.read(), exc.headers)
except Exception: # pragma: no cover - body is best-effort context
pass
# A 5xx can be a whole HTML page, and this text reaches the alert
# and the run log's error_summary.
if len(body) > ERROR_BODY_CHARS:
body = body[:ERROR_BODY_CHARS] + "..."
if exc.code == 401:
raise AuthError(f"401 rejected key {self.masked_key}: {body}") from exc
if exc.code == 404:
Expand Down
Loading
Loading