Repository navigation
feat(ticketing): add a watchdog for the /notify pipeline going silent - #203
Merged
Merged
Conversation
The alert-judgment/correlation pipeline (correlator.judge() and everything downstream of it) only ever runs once an alert actually arrives at /notify. There was no check for the inverse: n8n's upstream polling job can stop calling /notify entirely, fleet-wide, and nothing in this repo would ever notice -- confirmed this actually happened, 5 days of total silence across every grid, only caught by a customer asking the bot directly. Adds a scheduled watchdog (registered alongside the existing escalation Jira sweep) that checks four times a day whether notify_alert_deliveries has a row for ANY grid within NOTIFY_WATCHDOG_HOURS (default 24h). If not: - Files a Jira ticket (deliberately grid-less -- this is a statement about the pipeline, not any one site), deduped by searching for an existing open ticket with a fixed label first so it doesn't re-file every check while the same silence continues. - Posts a Telegram alert to ESCALATION_TELEGRAM_CHAT_ID. NOTIFY_WATCHDOG_ENABLED (default true) is the deliberate off switch for a deployment that never wires up /notify at all -- matching the same default-on convention every other scheduled batch in this deployment already uses (episodic distillation, the Grafana indexer, the escalation sweep). A CronTrigger at fixed wall-clock hours rather than an interval timer on purpose: this deployment redeploys multiple times a day, and an interval timer resets on every restart. New: notify_watchdog.py (the check + fire logic, dependency-injected for testability), NotifyAlertDeliveryRepository.most_recent_delivery_at() (fleet-wide, unlike the existing per-grid queries), and JiraTicketBackend.find_open_ticket_by_label() (the dedup search, scoped to open tickets only so a long-resolved prior ticket is never mistaken for "still tracked"). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Answers the user's question from the same incident investigation as #201/#202: "the fact that the alerts have been silent should flag a general O&M ticket, can the bot do that in case this happens again?"
The alert-judgment/correlation pipeline (
correlator.judge()and everything downstream of it) only ever runs once an alert actually arrives at/notify. There was no check for the inverse — n8n's upstream polling job can stop calling/notifyentirely, fleet-wide, and nothing in this repo would ever notice. Confirmed this actually happened: 5 days of total silence across every grid, only caught because a customer happened to ask the bot directly.I considered (and the user asked directly) whether episodic distillation could cover this instead — it can't: it's per-grid/per-org scoped (keyed by name-matching chat content), purely a passive read-time summary with no way to file a ticket or post a message, and (per #201's investigation) is itself currently silently broken. Piggybacking a reliability check on an already-unreliable, unrelated system would be the wrong foundation. This is a dedicated, deterministic check instead.
Change
A scheduled watchdog, registered alongside the existing escalation Jira sweep, checking four times a day whether
notify_alert_deliverieshas a row for any grid withinNOTIFY_WATCHDOG_HOURS(default 24h). If not:ESCALATION_TELEGRAM_CHAT_ID.NOTIFY_WATCHDOG_ENABLED(default true) is the admin-configurable off switch for a deployment that never wires up/notifyat all — matching the exact default-on convention every other scheduled batch in this deployment already uses (episodic distillation, the Grafana indexer, the escalation sweep itself). As a second line of defense, a table that has never had a single row (a fresh deployment, or one that genuinely never uses/notify) is treated as "not applicable," not silence — so an operator who forgets to flip the flag is still unlikely to get a false alarm.A
CronTriggerat fixed wall-clock hours (01:15/07:15/13:15/19:15 UTC), not an interval timer, on purpose: this deployment redeploys multiple times a day, and an interval timer resets on every restart — it could go long stretches never reaching its own interval.New:
notify_watchdog.py— the check + fire logic, fully dependency-injected for testability (no real Jira/Telegram/DB needed in tests).NotifyAlertDeliveryRepository.most_recent_delivery_at()— fleet-wide (unlike the existing per-gridlatest_downtime_sent_at), returns the rawsent_atstring matching that method's own established convention.JiraTicketBackend.find_open_ticket_by_label()— the dedup search, scoped to open tickets only (statusCategory != Done) so a long-resolved prior ticket is never mistaken for "still tracked."Test plan
test_notify_watchdog.py— pure-function coverage forsilence_detected/_parse_sent_at/env-var gating, plus fullrun_notify_watchdogorchestration coverage with injected fakes: disabled, not-silent, never-had-a-delivery, silent-and-files, already-tracked-skips, dedup-check-failure-still-files, ticket-creation-failure-still-alerts, telegram-failure-doesn't-raise, empty-chat-id-skips-gracefully, delivery-lookup-failure-returns-cleanly.test_notify_alert_delivery_repository.pyformost_recent_delivery_at(fleet-wide scope, empty table, ledger outage fail-open).test_jira_backend.pyforfind_open_ticket_by_label(found/not-found/no-credentials/error-status, and a direct JQL-string assertion that resolved tickets are excluded).chat_orchestratorsuite (uv run pytest tests/ -q, matching CI): 2524 passed, 1 pre-existing unrelated failure (confirmed identical onmainbefore this change).uv run pytest ../shared -q -n auto: pre-existing failures only, both confirmed unrelated.pre-commit run --all-filesclean./notifysilence).🤖 Generated with Claude Code