Skip to content

feat(ticketing): add a watchdog for the /notify pipeline going silent - #203

Merged
1vav merged 1 commit into
mainfrom
feat/notify-pipeline-watchdog
Sep 18, 2026
Merged

1vav merged 1 commit into
mainfrom
feat/notify-pipeline-watchdog

Conversation

@1vav

@1vav 1vav commented Sep 18, 2026

Copy link
Copy Markdown
Contributor

Summary

Answers the user's question from the same incident investigation as #201/#202: "the fact that the alerts have been silent should flag a general O&M ticket, can the bot do that in case this happens again?"

The alert-judgment/correlation pipeline (correlator.judge() and everything downstream of it) only ever runs once an alert actually arrives at /notify. There was no check for the inverse — n8n's upstream polling job can stop calling /notify entirely, fleet-wide, and nothing in this repo would ever notice. Confirmed this actually happened: 5 days of total silence across every grid, only caught because a customer happened to ask the bot directly.

I considered (and the user asked directly) whether episodic distillation could cover this instead — it can't: it's per-grid/per-org scoped (keyed by name-matching chat content), purely a passive read-time summary with no way to file a ticket or post a message, and (per #201's investigation) is itself currently silently broken. Piggybacking a reliability check on an already-unreliable, unrelated system would be the wrong foundation. This is a dedicated, deterministic check instead.

Change

A scheduled watchdog, registered alongside the existing escalation Jira sweep, checking four times a day whether notify_alert_deliveries has a row for any grid within NOTIFY_WATCHDOG_HOURS (default 24h). If not:

  • Files a Jira ticket — deliberately grid-less (this is a statement about the pipeline, not any one site), deduped by searching for an existing open ticket with a fixed label first so it doesn't re-file on every check while the same silence continues (a resolved ticket "forgets" and a fresh one files if it recurs later).
  • Posts a Telegram alert to ESCALATION_TELEGRAM_CHAT_ID.

NOTIFY_WATCHDOG_ENABLED (default true) is the admin-configurable off switch for a deployment that never wires up /notify at all — matching the exact default-on convention every other scheduled batch in this deployment already uses (episodic distillation, the Grafana indexer, the escalation sweep itself). As a second line of defense, a table that has never had a single row (a fresh deployment, or one that genuinely never uses /notify) is treated as "not applicable," not silence — so an operator who forgets to flip the flag is still unlikely to get a false alarm.

A CronTrigger at fixed wall-clock hours (01:15/07:15/13:15/19:15 UTC), not an interval timer, on purpose: this deployment redeploys multiple times a day, and an interval timer resets on every restart — it could go long stretches never reaching its own interval.

New:

  • notify_watchdog.py — the check + fire logic, fully dependency-injected for testability (no real Jira/Telegram/DB needed in tests).
  • NotifyAlertDeliveryRepository.most_recent_delivery_at() — fleet-wide (unlike the existing per-grid latest_downtime_sent_at), returns the raw sent_at string matching that method's own established convention.
  • JiraTicketBackend.find_open_ticket_by_label() — the dedup search, scoped to open tickets only (statusCategory != Done) so a long-resolved prior ticket is never mistaken for "still tracked."

Test plan

  • New test_notify_watchdog.py — pure-function coverage for silence_detected/_parse_sent_at/env-var gating, plus full run_notify_watchdog orchestration coverage with injected fakes: disabled, not-silent, never-had-a-delivery, silent-and-files, already-tracked-skips, dedup-check-failure-still-files, ticket-creation-failure-still-alerts, telegram-failure-doesn't-raise, empty-chat-id-skips-gracefully, delivery-lookup-failure-returns-cleanly.
  • New tests in test_notify_alert_delivery_repository.py for most_recent_delivery_at (fleet-wide scope, empty table, ledger outage fail-open).
  • New tests in test_jira_backend.py for find_open_ticket_by_label (found/not-found/no-credentials/error-status, and a direct JQL-string assertion that resolved tickets are excluded).
  • Full chat_orchestrator suite (uv run pytest tests/ -q, matching CI): 2524 passed, 1 pre-existing unrelated failure (confirmed identical on main before this change).
  • uv run pytest ../shared -q -n auto: pre-existing failures only, both confirmed unrelated.
  • pre-commit run --all-files clean.
  • Not yet verified against a real fired ticket/Telegram post in prod (would require actually inducing a /notify silence).

🤖 Generated with Claude Code

The alert-judgment/correlation pipeline (correlator.judge() and everything
downstream of it) only ever runs once an alert actually arrives at
/notify. There was no check for the inverse: n8n's upstream polling job can
stop calling /notify entirely, fleet-wide, and nothing in this repo would
ever notice -- confirmed this actually happened, 5 days of total silence
across every grid, only caught by a customer asking the bot directly.

Adds a scheduled watchdog (registered alongside the existing escalation
Jira sweep) that checks four times a day whether notify_alert_deliveries
has a row for ANY grid within NOTIFY_WATCHDOG_HOURS (default 24h). If not:
- Files a Jira ticket (deliberately grid-less -- this is a statement about
  the pipeline, not any one site), deduped by searching for an existing
  open ticket with a fixed label first so it doesn't re-file every check
  while the same silence continues.
- Posts a Telegram alert to ESCALATION_TELEGRAM_CHAT_ID.

NOTIFY_WATCHDOG_ENABLED (default true) is the deliberate off switch for a
deployment that never wires up /notify at all -- matching the same
default-on convention every other scheduled batch in this deployment
already uses (episodic distillation, the Grafana indexer, the escalation
sweep). A CronTrigger at fixed wall-clock hours rather than an interval
timer on purpose: this deployment redeploys multiple times a day, and an
interval timer resets on every restart.

New: notify_watchdog.py (the check + fire logic, dependency-injected for
testability), NotifyAlertDeliveryRepository.most_recent_delivery_at()
(fleet-wide, unlike the existing per-grid queries), and
JiraTicketBackend.find_open_ticket_by_label() (the dedup search, scoped to
open tickets only so a long-resolved prior ticket is never mistaken for
"still tracked").

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@1vav
1vav merged commit 6bf411d into main Sep 18, 2026
7 checks passed
@1vav
1vav deleted the feat/notify-pipeline-watchdog branch September 18, 2026 09:36
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant