Skip to content

fix(ticketing): wire alarm and grid-state facts into alert judgment - #200

Merged
1vav merged 1 commit into
mainfrom
fix/alert-judgment-context-gaps
Sep 12, 2026
Merged

1vav merged 1 commit into
mainfrom
fix/alert-judgment-context-gaps

Conversation

@1vav

@1vav 1vav commented Sep 12, 2026

Copy link
Copy Markdown
Contributor

Summary

The live alert-correlation LLM call writes every alert's "Likely cause" and
"Suggested action" lines and decides whether a new alert joins an existing
ticket. It never saw VRM alarm data or grid-state duration facts, even
though the prompt it's given explicitly instructs it to use both. Both now
reach the prompt.

These ship together because they're the same mistake in two places: when
the correlation prompt was rewritten around a new typed context object
(AlertJudgmentContext), two sources the old prompt fed the model had no
matching field in the new schema and were silently dropped rather than
erroring.

What was missing

  • VRM alarm / power-trend evidence. The fetch is real (VRM's actual
    alarm-log and diagnostics endpoints, scoped to the right site and a
    30-minute window) and the raw data reaches the context-assembly layer —
    but AlertTelemetry, the schema that shapes what the prompt actually
    sees, never declared the three fields, so pydantic's extra="ignore"
    dropped them before the LLM call. The prompt names "overload" as the
    literal example of what this evidence should surface.
  • Grid-state duration facts (is_hps_on / is_hps_on_updated_at / DCU
    status roll-up). The prompt's "Root Cause Rules" section — the mechanism
    for deciding whether a new alert is a child of an existing grid-off state
    — depends on this, and the pre-rewrite prompt path fetched it. The
    rewritten path never called the fetch at all.

Fix

  • Add the three alarm/power-trend fields to AlertTelemetry, matching the
    names the prompt already expects, with a small alias step mapping the
    raw provider's own key names onto them before validation — so neither
    the provider nor the prompt vocabulary has to change, only the one seam
    between them.
  • Add a grid_operational_facts source to AlertJudgmentContext, gathered
    by the assembler the same resilient way every other source already is
    (bounded timeout, failure isolation, availability tracking). Left as a
    raw dict rather than a new narrow sub-model, since a stricter schema is
    exactly what dropped the alarm fields the first time.
  • Wire the new provider at the one real call site and add the
    corresponding prompt section.

Testing

Extended the assembler and prompt-builder unit tests to pin both sources
surviving end to end — raw provider data in, correct field names in the
built prompt out — plus the EMPTY-vs-FAILED distinction for the new
grid-facts source (a lookup miss is "no extra context," not degradation).
Full chat_orchestrator suite and pre-commit run --all-files clean.


Compound Engineering
Claude Code

The live alert-correlation LLM call (AlertCorrelator.judge /
_build_judgment_prompt, the v2 path that's been the only one running in
prod since ALERT_LLM_JUDGMENT_ENABLED defaulted on) was silently missing
two context sources that ticketing.correlation.prompt's own text already
instructs the model to use:

1. VRM alarm/power-trend evidence. get_recent_production_errors_and_power
   genuinely fetches this from VRM's real alarm-log/diagnostics endpoints
   and UrgentAlertContext.telemetry() carries it in the raw dict handed to
   AlertJudgmentContextAssembler -- but AlertTelemetry never declared the
   three fields, so they were dropped at the pydantic-validation boundary
   (extra="ignore") before the prompt was ever built. The prompt names
   "overload" as the literal example of what this evidence should surface.

2. Grid-state duration facts (is_hps_on / is_hps_on_updated_at / DCU status
   roll-up). The legacy prompt builder (_build_prompt / _call_llm, used by
   the pre-v2 decide() path) fetches these via
   correlation_rules.get_grid_operational_context; the v2 judge() path
   never called it at all. The prompt's "Root Cause Rules" section -- the
   primary mechanism for deciding whether a new alert is a child of an
   existing grid-off/isolated state -- depends entirely on this duration
   signal, which the live judgment call has had zero access to since the
   v2 path went live.

Both were lost the same way: the v2 prompt/schema was written fresh
instead of by diffing against everything the legacy prompt fed the model,
so two sources with no equivalent field in the new typed models were
silently dropped rather than erroring.

Fix:
- Add active_ve_bus_errors / recent_ve_bus_errors_30min /
  power_trend_past_30min to AlertTelemetry, matching the names the prompt
  already expects. A small alias step maps the raw telemetry provider's own
  key names (active_alarms / recent_alarms_30min / power_history_30min)
  onto these before validation, so neither the provider nor the prompt
  vocabulary has to change -- only the one seam between them.
- Add a grid_operational_facts source to AlertJudgmentContext, gathered by
  the assembler the same resilient way every other source is (bounded
  timeout, failure isolation, availability tracking). Left as a raw dict
  rather than a new narrow sub-model -- a stricter schema is exactly what
  dropped the alarm fields the first time.
- Wire a grid_operational_facts_provider at the one real call site
  (app.py's judgment-context assembly) and add the corresponding section to
  _build_judgment_prompt.

Tests: extended the assembler and prompt-builder unit tests to pin both
fields surviving end to end (raw provider -> typed context -> built
prompt), plus the empty/EMPTY-vs-FAILED distinction for the new source.
Full chat_orchestrator suite passes (2490 passed; the 4 test_loop_detector
failures are the known .env.example LOOP_DETECTION_THRESHOLD mismatch,
unrelated to this change and reproducible on main). pre-commit run
--all-files clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@1vav
1vav merged commit fa14aab into main Sep 12, 2026
7 checks passed
@1vav
1vav deleted the fix/alert-judgment-context-gaps branch September 12, 2026 11:53
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant