Skip to content

My Mesh: show estimated zero-hop and flood advert intervals on repeater cards - #352

Open
dborup wants to merge 5 commits into
masterfrom
codex/my-mesh-advert-cadence-304
Open

dborup wants to merge 5 commits into
masterfrom
codex/my-mesh-advert-cadence-304

Conversation

@dborup

@dborup dborup commented Oct 7, 2026 •

Copy link
Copy Markdown
Owner

Summary

Add separate estimated zero-hop and flood advert intervals to repeater cards on My Mesh (part of #304).

This reuses the existing #245 estimator and its cached, route-aware advert scan. The My Mesh card opts in on its existing node-health request (/api/nodes/{pk}/health?include=advertIntervals). It does not make another HTTP request per card or duplicate the interval algorithm. For a claimed node, the health panel (card click / Full health) reuses the same URL through one healthPath() helper, so the client cache answers it. Nodes outside My Mesh keep the plain /health URL. Non-repeater and default health responses do not pay for the advert lookup. Hidden identities fail closed before the lookup.

The card distinguishes estimated, too few, none observed, irregular, and unavailable states. It shows the sample count and latest advert when present, and says the values are observed estimates from at most the newest 20 adverts per class, not device configuration. Missed receptions may skew the estimate. While route-mask backfill is not complete (or its status is missing), the card marks the route classes provisional and forwards the existing typed advertRouteBackfill status. The full node page remains the detailed view.

Base

The base is master. #351 (activity chart and compact observer list) merged as 1f423e06. This branch now merges origin/master with a normal merge commit (no rebase or force-push). The merge keeps #351's rounds 2–3:

  • clipped observer names stay recoverable on every card (title, plus Full names → revealed by revealClippedObserverButtons);
  • the fraction-width-safe 24h pre-filter;
  • the escapeAttr quote/XSS regression in test-issue-304-my-mesh-e2e.js.

#352's cadence block renders after the observer row. The E2E fixture keeps #351's eight mixed-role cards and adds the cadence data to repeaters 0–2.

Validation

  • cmd/server: go test -race ./... passed. cmd/ingestor: go test ./... passed.
  • sh test-all.sh (234/234) and node test-frontend-helpers.js (712) passed. One unrelated timing flake in test-channels-client-state-152.js (unchanged from master) was seen once under load and passed on re-run.
  • test-issue-304-my-mesh-e2e.js (desktop 1280 and mobile 375), test-issue-2027-my-mesh-node-page-e2e.js (25/25) and test-home-coverage-e2e.js (12/12, 5 runs) passed with a local Go server on the CI-prepared e2e-fixture.db.
  • On that server, /health?include=advertIntervals and /nodes/{pk}?include=advertRoutes return identical intervals and backfill status for the feat(nodes): estimated advert intervals (flood / zero-hop) per node, then mesh-wide settings analytics #245 seed node. Default /health returns neither field.
  • XSS preflight (--diff origin/master and --file public/home.js), node --check, and git diff --check passed.

Performance and privacy

The existing nodeAdvertRoutes cache has a 30-second TTL, per-node invalidation, a 128-entry cap, and singleflight. A backend test checks that three opt-in health calls perform only one scan, and that default and non-repeater calls perform none. A cold scan still costs more than ordinary health, so the field is opt-in. The same identity-hidden gate as node detail runs before the cached breakdown is accessed.

No merge or deployment is requested here.

🤖 Generated with Claude Code

@dborup dborup left a comment

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review against codex/my-mesh-activity-304:

[P2] Preserve the route-backfill uncertainty on the My Mesh card. The existing node-detail include=advertRoutes response includes advertCounts.route_mask_backfill and node-adverts.js warns when its status is not complete. This PR forwards only advertIntervals to health. Until route-mask backfill completes, legacy adverts can be classified from their first-stored route, so a normal-looking zero-hop/flood cadence on the card can be provisional. Please carry the existing backfill status through the opt-in health response and show the same caveat (or suppress the estimate) when it is not complete. Add a backend/browser regression for that state.

The rest of the approach looks sound: it reuses #245 rather than duplicating interval logic, keeps the cold scan opt-in and cached, and applies the identity-hidden gate before the lookup. I ran the full server, ingestor and frontend suites plus the My Mesh Chromium tests; all passed after the existing Kpa-clawbot#2027 fixture was updated.

@dborup

dborup commented Oct 7, 2026

Copy link
Copy Markdown
Owner Author

Review feedback addressed (commit 1f54e5a4):

  1. Opted-in repeater health now includes the existing typed advertRouteBackfill status behind the same privacy gate as advertIntervals.
  2. My Mesh labels the route classes provisional whenever that status is not complete, including when it is missing from an older response. The tooltip explains why the flood/zero-hop estimate may change.
  3. Added backend tests for pending/complete status, default/non-repeater/hidden suppression and cache reuse, plus desktop/mobile browser assertions for pending, complete and missing states. Updated the API documentation.
  4. Re-ran the full server suite, 227-file frontend suite, both My Mesh Chromium tests and XSS preflight; all passed. The unchanged ingestor suite had already passed after running outside the local-port sandbox.

@dborup dborup left a comment

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Follow-up review of commit 1f54e5a4 against the stacked base: the route-backfill finding is resolved. The opt-in health response now forwards the existing typed backfill status behind the identity-hidden gate; the card warns for pending, backfilling, and unknown status, while a complete status stays quiet. Backend and desktop/mobile browser regressions cover the distinction. I found no further actionable issues in this follow-up.

Local validation is green (full Go server and ingestor suites, 227 frontend test files, both My Mesh Chromium tests). This stacked PR has no automatic PR check because the workflow filters pull requests to master; a manual branch run is being used for CI validation. No merge or deployment is implied.

@dborup

dborup commented Oct 7, 2026

Copy link
Copy Markdown
Owner Author

Taking this PR over for an independent review. It is stacked on #351; it will be retargeted to master after #351 merges.

@dborup-agent

Copy link
Copy Markdown
Collaborator

Review — CS-MacBook PR#352 — head 1f54e5a

Dom: APPROVE with nits

Independent, read-only review of this PR's own diff on top of its stacked base codex/my-mesh-activity-304 (git diff origin/codex/my-mesh-activity-304...1f54e5a4). #351 is reviewed separately and is out of scope here. Evidence is tagged [T] test/live-run, [A] architecture/source-read, [K] code-read.

Findings

# Sev Finding
1 nit Route classes provisional renders whenever advertIntervals is present and backfill status ≠ complete, even when both classes have no estimate (none_observed / unavailable). Cosmetically odd — there is no classification to be provisional about — but harmless and arguably still honest.
2 nit The frontend sends ?include=advertIntervals for every My Mesh node, including non-repeaters. The server correctly skips the scan for non-repeaters and the card gates rendering on isRepeater, so there is no extra cost; the param is just slightly broader than strictly needed.
3 info CI "✅ Go Build & Test" is red on TestServerProcess_WaitsForTheSchemaThenStartsOnce (SIGTERM → exit -1), a process/startup-timing test unrelated to this diff. It passes locally on the merged tree with -race (see Tests). Treated as an environmental flake, not a regression from this PR.

No correctness, privacy, or XSS issues found.

Answers to the review points

1. Separate zero-hop & flood intervals, reuse of #245/#247/#292 (DRY), values vs node detail.
Confirmed reuse, no duplication. handleNodeHealth opts in via s.nodeAdvertRoutes(pubkey, now) — the exact same per-pubkey cached function and estimator that node detail uses (routes.go:2039 → res.intervals for node detail; routes.go:2118 → adverts.intervals for health). Both assign the identical res.intervals, so values are identical by construction, from one cache entry. [A][K] The card shows zero_hop and flood as separate rows with sample count and last-advert age. The include parser is factored into a shared wantsNodeInclude, matching the existing advertRoutes semantics (repeated + comma-separated). [K]

2. Provisional route-classification warning — shown in the right cases?
Yes. Condition is intervals && backfill?.status !== 'complete'. Verified live at both widths: pending → shown, missing/undefined backfill → shown, complete → suppressed. The backend forwards the existing typed advertRouteBackfill (RouteMaskBackfillStatus) read fresh, behind the same identity-hidden gate as the estimate; when intervals are emitted the status is always emitted alongside, and the frontend still defensively treats an absent status as provisional (older-server case). [T][A]

3. cmd/server read-only, no new untyped maps, hardcoded colours, XSS.
cmd/server stays read-only — no INSERT/UPDATE/DELETE/Exec/Prepare added; the scan is in-memory cached reads only. [K] The only map[string]interface{} in added non-test code is a read type-assertion on the existing GetNodeHealth result map (map[string]interface{} already), not a new structure. [K] No hardcoded colours — new CSS uses var(--status-yellow) / var(--text-muted). [K] XSS gate clean (check-xss-sinks.sh --file on the merged home.js, exit 0); all interpolated values are fixed labels, numbers, or timeAgo() output — no node-controlled field reaches innerHTML unescaped. [T][K]

4. Fork-guards / closing keywords / commit author.
PR touches no .github/ workflow, so fork-guards are unchanged (N/A). No closing keywords in the PR body or either commit (body explicitly says "No merge or deployment is requested"). Both commits authored by dborup <kontakt@meshview.dk>. [K]

Acceptance criteria → tests

  • Opt-in only, repeater-only, hidden fails closed, cache reuse (1 scan for 3 lookups), pending/complete forwarding, privacy-gates-before-cache: cmd/server/node_health_advert_intervals_test.go (TestNodeHealthAdvertIntervalsOptInRoleAndPrivacy, ...FailsClosedOnIdentityLookupError). [T]
  • Separate intervals rendered, provisional for pending/missing and quiet for complete, "insufficient"/"unavailable" states, long-name/XSS observer handling, compact height: test-issue-304-my-mesh-e2e.js (desktop 1280 + mobile 390). [T]
  • Opt-in URL does not disturb node-page nav / full-health fetch, no console errors: test-issue-2027-my-mesh-node-page-e2e.js (25/25). [T]

Tests & mutants (merged tree git merge-tree origin/master 1f54e5a4 = 3ec8823c, over origin/master 987c220d)

  • Red-before / green-after [T]: removing the handleNodeHealth opt-in hunk → TestNodeHealthAdvertIntervalsOptInRoleAndPrivacy fails (repeater missing intervals); restored → ok.
  • Mutant 1 [T]: invert privacy gate (!hidden → hidden) → caught (repeater missing intervals).
  • Mutant 2 [T]: role gate repeater → companion → caught (non-repeater exposes intervals).
  • Full server suite go test -race ./... (cmd/server) [T]: ok github.com/corescope/server 542.9s, no FAIL — includes the CI-flaky TestServerProcess_WaitsForTheSchemaThenStartsOnce, green here.
  • sh test-all.sh [T]: 232 passed, 0 failed.
  • node test-frontend-helpers.js [T]: 709 passed, 0 failed.
  • Browser [T]: rendered real home.js with stubbed health at 1280×900 and 375×812 — 4 cards, correct cadence (≈ 2 h, not enough adverts, not observed, irregular, unavailable), 2 provisional badges on pending + missing-backfill cards, none on the complete card, compact heights (≤317 px), no console or page errors.

Not verified

  • Go cmd/ingestor suite — unchanged by this diff (not touched), not re-run.
  • A live end-to-end /health?include=advertIntervals vs /nodes/{pk}?include=advertRoutes comparison against a seeded e2e-fixture.db Go server was not run; value identity is established at source level instead (same function, same cache).
  • --diff origin/master XSS mode (which needs the PR as git HEAD) was substituted with the stronger whole-file --file scan on the merged home.js.
  • My Mesh: accurate 24h node activity and compact repeater observer cards #351's own diff (reviewed separately).

@dborup
dborup changed the base branch from codex/my-mesh-activity-304 to master October 8, 2026 07:28
dborup and others added 3 commits October 8, 2026 07:30
Bring #351 rounds 2-3 (merged as 1f423e0) under #352's advert cadence.

Conflicts:
- public/home.js: keep master's clipped-name disclosure (title +
  revealClippedObserverButtons) and add #352's cadence block after the
  observer row.
- test-issue-304-my-mesh-e2e.js: keep master's eight mixed-role cards,
  wide/clipped names and the escapeAttr quote test; carry #352's
  advertIntervals/advertRouteBackfill fixtures on repeaters 0-2 and the
  opt-in health URL. Also pin that the no-estimate repeater shows no
  provisional badge and a non-repeater shows no cadence.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…che (#352)

#352 moved the card's health request to ?include=advertIntervals but
loadHealth still fetched the plain URL, so a card click became a second
network request instead of a client-cache hit. CI's home-coverage E2E
caught it as a race (the panel was still "Loading…").

Pin that a card click renders from the card's own URL, and that a node
outside My Mesh keeps the plain URL (no advert scan).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…352)

One healthPath() helper builds the card request and the panel request.
For a claimed node the panel asks for the same ?include=advertIntervals
URL the card already fetched, so the client cache answers it as before
#352. Nodes outside My Mesh keep the plain URL and skip the advert scan.
Update the Kpa-clawbot#2027 harness comment that still said Full health used the
plain URL.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@dborup-agent

Copy link
Copy Markdown
Collaborator

Rapport — CS-pve-agent3 PR#352 master-sync — head a5c782a

Review feedback addressed (commit a5c782ae; merge 3592c115, test 7b37e1b6, fix a5c782ae)

  1. Base sync, My Mesh: accurate 24h node activity and compact repeater observer cards #351 N5 (merge, not rebase). I merged origin/master (1f423e06) into the branch with an ordinary merge commit, 3592c115. There was no rebase, amend or force-push, so 1f54e5a4 stays an ancestor. Two files conflicted. I resolved both by keeping both sides:
  2. New, found by CI on the merged head: a card click no longer hit the client cache (blocking, fixed). My Mesh: show estimated zero-hop and flood advert intervals on repeater cards #352 moved the card request to /health?include=advertIntervals, but loadHealth() (card click / Full health) still fetched plain /health. Before My Mesh: show estimated zero-hop and flood advert intervals on repeater cards #352 that panel was a client-cache hit on the card's response. After My Mesh: show estimated zero-hop and flood advert intervals on repeater cards #352 it is a cache miss and a second network request. test-home-coverage-e2e.js ("clicking a My Mesh card loads health detail panel: expected .health-banner after card click") failed in CI run 37745490854 because the panel was still "Loading…". I reproduced it locally in 1 of 3 runs. My Mesh: show estimated zero-hop and flood advert intervals on repeater cards #352's own earlier CI never got past the Go job, so this E2E had not run on the PR before. The feat(home): link My Mesh node cards to the node page Kpa-clawbot/CoreScope#2027 harness registers both URLs, which hid the split [T][K].
  3. MacBook nit 1 (provisional shown when neither class has an estimate): not changed. I disagree that it is purely cosmetic. classifyAdvertRouteRow puts an advert with no backfilled route_mask and a NULL route_type in unknown, and the estimator does not use that bucket. Once backfill writes its mask, the same advert can become zero_hop or flood. So none_observed / too_few / irregular can also change until backfill is complete. The badge only renders when advertIntervals is present, and the new card-3 assertion pins that [A][K].
  4. MacBook nit 2 (include=advertIntervals sent for non-repeaters): left as is, not trivial. The card learns the role from the same health response (h.node.role). My Mesh's stored entries hold only {pubkey, name}, so limiting the opt-in to repeaters would need a prior request or a new client-side role cache. The server already skips the scan for non-repeaters, and a backend test pins that (0 scans), so this costs nothing [A][K].
  5. MacBook item 3 (CI Go flake): informational. See CI below.
  6. PR description: I updated it for the master base. I removed the stacked-on-My Mesh: accurate 24h node activity and compact repeater observer cards #351 text, added a "Base" section that lists what the merge preserves, and refreshed the validation section.

Tests

  • [T] cmd/server: go test -race -count=1 ./... gives ok github.com/corescope/server 640.6s.
  • [T] cmd/ingestor: go test -count=1 ./... gives ok github.com/corescope/ingestor 208.2s. internal/regions passes too.
  • [T] sh test-all.sh on a5c782ae: 234 passed, 0 failed. On 3592c115 it was 233 passed, 1 failed (234 files). The failure was test-channels-client-state-152.js: R4-3 S2: a remap before computeChannelHash resolves restarts loading: pane must not hang on "Decrypting…". The test and public/channels.js are byte-identical to master, and the test does not load home.js. It passed 5/5 standalone runs, then failed once in a 25-run loop while the race suite was loading the machine. This is a timing flake that is neither flaky: #180 packets URL E2E 'reload shows the same list' depends on fixture age (default 15-min window) #271 nor test(server): data race between startBackgroundIndexBuilds and PacketStore.Load under -race (TestHashMigrate_LogsMaxWriteLockHold_215) #301, and I found no existing issue for it. I did not open one.
  • [T] node test-frontend-helpers.js: 712 passed, 0 failed.
  • [T] XSS gate: check-xss-sinks.sh --diff origin/master rc=0, --file public/home.js rc=0. node --check and git diff --check origin/master HEAD pass.
  • [T] node test-frontend-helpers.js and the XSS gate (--diff origin/master rc=0, --file rc=0) were re-run on a5c782ae and passed. The Go suites below ran on 3592c115; the later commits touch only frontend files.
  • [T] Local Go server on e2e-fixture.db, prepared as in CI: freshen, the CI inline seed SQL, corescope-migrate, then the 2073/199/245 seeds.
    • CHROMIUM_REQUIRE=1 node test-issue-304-my-mesh-e2e.js passes at 1280×900 and 375×812 touch.
    • test-issue-2027-my-mesh-node-page-e2e.js passes 25/25.
    • For the feat(nodes): estimated advert intervals (flood / zero-hop) per node, then mesh-wide settings analytics #245 seed node, /health?include=advertIntervals and /nodes/{pk}?include=advertRoutes return identical advertIntervals (flood 43200 s, zero-hop 7200 s) and the same backfill status {pending, remaining:null}. Plain /health has neither field. This covers the MacBook review's "not verified" live comparison.
  • [T] Browser screenshot at 1280 of the merged card: cadence rows only on repeaters, Route classes provisional on the pending and missing-backfill cards, Full names → on clipped cards, no page errors.

Mutants (each applied alone, then reverted; tree clean afterwards)

# Guards Mutation Result
M1 #352 cadence kept through merge drop ${cadenceHtml} killed [T]: .mnc-advert-cadence not found (e2e:123)
M2 #351 revealClippedObserverButtons drop the call killed [T]: card 3 must offer the observer dialog
M3 #351 escapeAttr XSS (O9) title="${obsFullName(o)}" killed [T]: observer span gained attributes: class,onmouseover,title
M4 #352 provisional on missing status !== 'complete' changed to === 'pending' killed [T]: card 2 .mnc-advert-provisional not found (e2e:128)
M5 cadence repeater-only (new assertion) isRepeater ? changed to true ? killed [T]: non-repeater card must not show advert intervals
M6 no badge without intervals (new assertion) drop intervals && killed [T]: e2e:130
M7 #351 24h pre-filter < a.startSecond changed to < a.WindowStart killed [T]: TestNodeActivity24hPreFilterIgnoresFractionWidth
M8 #352 privacy gate !hidden changed to hidden killed [T]: TestNodeHealthAdvertIntervalsOptInRoleAndPrivacy: repeater missing intervals
M9 panel reuses card URL (new finding) loadHealth back to plain /health killed [T]: card click must render health from the card's request: Failed to load node health.
M10 unclaimed node keeps plain URL healthPath(pubkey, true) in loadHealth survived the first test version; I added the routes assertion, and now it is killed [T]: '/health?include=advertIntervals' ≠ '/health'

The merge has no new production logic, so M1–M8 serve as its red runs. The new finding has a real red-before-fix commit (7b37e1b6).

CI per job

Head a5c782ae, run 37750573038 (pull_request) [T]:

Job Result
✅ Go Build & Test pass (24m21s)
🎭 Playwright E2E Tests pass (22m7s)
🏗️ Build & Publish Docker Image pass (46s)
📦 Release Artifacts skipped
🚀 Deploy Staging skipped
📝 Publish Badges & Summary skipped

Head 3592c115 (merge only), run 37745490854: Go Build & Test passed. Playwright E2E failed on test-home-coverage-e2e.js, which is the finding in item 2, not a known flake. I fixed it instead of re-running. No job had to be re-run for #271 or #301.

Leftovers

  • test-channels-client-state-152.js R4-3 S2 is timing-flaky under load (above). It is unrelated to this PR and has no tracking issue.
  • MacBook nits 1–2 are not applied, for the reasons in items 3–4.
  • The PR is currently not a draft (isDraft: false, the state before this session). I did not change its state.

@dborup

dborup commented Oct 8, 2026 •

Copy link
Copy Markdown
Owner Author

Non-blocking review note on head a5c782ae: My Mesh now fetches the opted-in health response with the existing CLIENT_TTL.nodeHealth cache (240 seconds; public/home.js:307, public/app.js:218). The server-side advert-route cache is only 30 seconds and reads the backfill status fresh, but the browser may continue displaying an old interval or the “Route classes provisional” badge for up to four minutes after new adverts arrive or backfill completes. That may be an acceptable dashboard freshness trade-off; please confirm it is intentional, or use a shorter TTL for this opted-in response. I found no merge-blocking issue in the current PR diff.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants