Skip to content

test(e2e): keep slide-over packets rows inside the fixture's time window - #63

Merged
dborup merged 1 commit into
masterfrom
codex/fix-slideover-gate-time-window
Sep 17, 2026
Merged

dborup merged 1 commit into
masterfrom
codex/fix-slideover-gate-time-window

Conversation

@dborup

@dborup dborup commented Sep 17, 2026

Copy link
Copy Markdown
Owner

Problem

In the CI integration draft #61 (run 35189733496) the "Slide-over E2E flake-gate (Kpa-clawbot#1616, --repeat-each=3)" failed on its first repetition:

✗ packets@800: page renders + first row exists: page.waitForFunction: Timeout 30000ms exceeded.
✗ packets@800: clicking row opens slide-over with backdrop: click setup failed: {"ok":false,"why":"no row","rowCount":0}

The same test had passed 27/27 in the main E2E step about five minutes earlier.

Mechanism

  • The e2e job freshens test-fixtures/e2e-fixture.db once before starting the server (newest packet = "now"); no new packets arrive afterwards.
  • The packets page defaults to a 15-minute window (public/packets.js:759-761) and requests /api/packets?since=now-15min. At widths ≤1024px a stored window above 180 minutes is reset to 15.
  • test-slideover-1056-e2e.js did not set a window, so the packets scenario used the 15-minute default.
  • CI timeline: fixture freshened 06:41:11; main-step slide-over passed 06:52 (~11 min, rows still inside the window); flake gate started 06:57:02 (~16 min) — by then no fixture packet was inside the window, the API returned 200 with total: 0, the table showed only "No packets found", and tr[data-hash] could never appear.

This was reproduced locally before this change (isolated servers, CI-style fixture copies, same test and gate loop), identically on master and on the integration tree: 6/6 passes shortly after freshening; 6/6 failures with exactly the CI errors once more than 15 minutes had passed; a probe showed API packets decaying 55 → 26 → 16 → 5 → 1 → 0 over ~15 minutes, while a 60-minute window at the same age still returned 138–150 packets. Running the 68 E2E tests that sit between the main slide-over run and the gate, then the gate, within 15 minutes passed 3/3, so the earlier tests do not affect it — elapsed time does.

Change

  1. test-slideover-1056-e2e.js — for the packets scenario only, context.addInitScript sets meshcore-time-window to "180" before the first navigation, with a comment on the static fixture and the ≤1024px limit. Assertions, selectors, timeouts, the other scenarios and the gate's repeat count are unchanged.
  2. .github/workflows/deploy.yml — timeout-minutes: 150 on the e2e-test job only.

Why both: 180 minutes is the largest window packets.js keeps at the test's 800px viewport, so the test alone cannot cover an arbitrarily long job. The job previously had no timeout (GitHub default 360 minutes). With a 150-minute job limit, the time from fixture freshening (inside the job) to the last gate run is always below the 180-minute data window; a job that runs that long now fails with an explicit timeout instead of a misleading "no row". The only observed e2e job run took ~18 minutes.

Semantic YAML comparison against master: the single difference is e2e-test.timeout-minutes: nil -> 150. Triggers, permissions, needs, fork-guards and deploy/publish conditions are unchanged; the cmd/server fork-guard workflow tests pass.

Regression evidence (local, before/after)

Isolated local Go server per fixture copy (prepared like CI: freshen + seed + migrate). The "aged" copy had every relevant timestamp (transmissions, observations, nodes, observers, neighbor edges) shifted back exactly 20 minutes before server start — no waiting; relative gaps verified identical to the fresh copy.

Fixture Probe (default window) Probe (180) Old test New test Gate: new test ×3, set -e -o pipefail
Fresh 200, 54 packets, window 15 200, 499 packets, window 180 27/27 27/27 3/3 pass
Aged 20 min 200, 0 packets, "No packets found" 200, 499 packets, fTimeWindow=180 24/27 — fails on packets@800: … first row exists (same 3 failures as CI) 27/27 3/3 pass

node --check, git diff --check, YAML parse: clean.

Limitations

  • Local browser: installed Google Chrome 151 with DNS restricted to localhost (CI uses Playwright's bundled Chromium with network access); public/ rather than CI's public-instrumented/. Leaflet cannot load offline (L is not defined on every run, passing and failing alike); it does not affect the packets table.
  • Other E2E tests that open the packets page without an explicit window could have the same latent dependency on elapsed time; they were not audited here.
  • GitHub CI for this change is being exercised through the CI integration draft CI integration only — do not merge (#56 + #59 + #60 + #62 + #63) #61.

🤖 Generated with Claude Code

… window

The e2e fixture is freshened once before the server starts, and the
packets page defaults to a 15-minute window. The slide-over flake gate
runs after the main E2E step, ~16 minutes in, when that window no longer
contains any fixture packet, so packets@800 waited 30s for a row that
could not appear. Pin the packets scenario to a 180-minute window, the
largest packets.js keeps at <=1024px, and bound the e2e job at 150
minutes so the gate always runs inside it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@dborup
dborup merged commit 693eb04 into master Sep 17, 2026
5 of 6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant