You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Coalesce timeline updates by poll run, not a 15-minute window #50
Raised by the code review of #48. Pre-existing, so it was left out of scope there.
What the code does
esb_site/model.py groups an outage's recorded changes into reader-visible updates with a time heuristic, COALESCE_WINDOW (15 minutes), in _build_updates and _envelope_updates. Changes within 15 minutes of the first change of a run fold into one update. #48 anchored the window to that first change so it no longer slides.
The window exists because a run's list response and detail fetch record their changes separately, seconds apart. Without it, a plain Fault -> Restored transition reads as two updates.
Why the heuristic can be wrong
It assumes a poll run lasts well under 15 minutes, and that different runs are always further apart than that. Neither always holds:
Storm runs are long. A run can take up to RUN_BUDGET_S (24 minutes; see notes/storms.md). A list change at T and the same outage's detail change at T+16 min then become two updates for one poll, which is the problem the window exists to prevent.
Hand-run polls are close together. Two separate runs 10 minutes apart fold into one update, and the earlier state disappears from the timeline. test_runs_close_together_do_not_slide_the_window in Fix seven bugs a retroactive review found in model.py #48 asserts this folding as current behaviour.
Proposal
The run table records started_at_utc and finished_at_utc (esb_outages/store.py). Grouping each change by the run whose interval contains its observed_at_utc would fix both directions at the root.
Things to check first:
Does every change's timestamp fall inside some run's [started_at, finished_at]? Early runs may be incomplete: 86 lack status (see the data-shape traps in CLAUDE.md). A change outside every run needs a fallback, probably the current window.
_envelope_updates works from the members' update times, not from raw changes, so it would need run membership carried on each Update.
_first_estimate already reads the raw change log, so it is unaffected.
Measure against esb-data how many updates, and which segment boundaries, change. If the change is worth making, add a dated entry to notes/grading.md § The update disclosure and update the CLAUDE.md Settled row "Coalescing changes within 15 minutes into one update".
Raised by the code review of #48. Pre-existing, so it was left out of scope there.
What the code does
esb_site/model.pygroups an outage's recorded changes into reader-visible updates with a time heuristic,COALESCE_WINDOW(15 minutes), in_build_updatesand_envelope_updates. Changes within 15 minutes of the first change of a run fold into one update. #48 anchored the window to that first change so it no longer slides.The window exists because a run's list response and detail fetch record their changes separately, seconds apart. Without it, a plain Fault -> Restored transition reads as two updates.
Why the heuristic can be wrong
It assumes a poll run lasts well under 15 minutes, and that different runs are always further apart than that. Neither always holds:
RUN_BUDGET_S(24 minutes; seenotes/storms.md). A list change at T and the same outage's detail change at T+16 min then become two updates for one poll, which is the problem the window exists to prevent.test_runs_close_together_do_not_slide_the_windowin Fix seven bugs a retroactive review found in model.py #48 asserts this folding as current behaviour.Proposal
The
runtable recordsstarted_at_utcandfinished_at_utc(esb_outages/store.py). Grouping each change by the run whose interval contains itsobserved_at_utcwould fix both directions at the root.Things to check first:
[started_at, finished_at]? Early runs may be incomplete: 86 lackstatus(see the data-shape traps in CLAUDE.md). A change outside every run needs a fallback, probably the current window._envelope_updatesworks from the members' update times, not from raw changes, so it would need run membership carried on eachUpdate._first_estimatealready reads the raw change log, so it is unaffected.Measure against esb-data how many updates, and which segment boundaries, change. If the change is worth making, add a dated entry to
notes/grading.md§ The update disclosure and update the CLAUDE.md Settled row "Coalescing changes within 15 minutes into one update".