Skip to content

test(ingestor): make neighbor metrics pruning deterministic - #33

Merged
adminopenclaw8-sketch merged 1 commit into
masterfrom
codex/fix-prune-neighbor-metrics-test-date
Sep 13, 2026
Merged

adminopenclaw8-sketch merged 1 commit into
masterfrom
codex/fix-prune-neighbor-metrics-test-date

Conversation

@dborup

@dborup dborup commented Sep 13, 2026 •

Copy link
Copy Markdown
Owner

Summary

Small, isolated prerequisite fix: makes TestPruneOldNeighborMetrics (cmd/ingestor) deterministic. This does not touch Go version, Dockerfiles, deploy.yml, production retention behavior, database migrations, or any other test — see "Scope audit" below.

Baseline: origin/master at mission start / this branch's parent commit — fda24ca575a9e77a3d9acce2aad249284e0d4f2f
Head: 1c6a8afca4ffcee85f95a11d598486fe782b921f

Relationship to #22: none. #22 (bump Dockerfile base images) is untouched — its branch, head SHA (207304ea0d93e80feb7e1d5a054caba8e639e63d), and files were re-verified unchanged at the end of this work. This PR does not approve, endorse, or merge #22 and carries no opinion on it. It exists only because independent research surfaced that TestPruneOldNeighborMetrics fails on both Go 1.22.12 and Go 1.27.1 — baseline nondeterminism, not a Go-version regression — and a Go-only PR should not have to carry an unrelated, already-broken test as noise in its CI run.

Root cause

TestPruneOldNeighborMetrics hardcoded its "recent" fixture to a fixed calendar date:

recent := "2026-07-26T12:00:00Z"
...
if n, err := store.PruneOldNeighborMetrics(30); ...

PruneOldNeighborMetrics(30) computes cutoff := time.Now().UTC().AddDate(0, 0, -30) and deletes rows with timestamp < cutoff. Once real time passed 2026-08-25 (30 days after the fixture date), the "recent" row itself crossed the retention boundary and started getting pruned alongside the intentionally-old row, failing the test's own assertion.

Before/after retention arithmetic (reproduced against unmodified fda24ca5, run at 2026-09-13T09:28:15Z)

value
now (test run time) 2026-09-13T09:28:15Z
cutoff = now − 30d 2026-08-14T09:28:15Z
fixture old 2020-01-01T00:00:00Z → before cutoff → pruned (correct)
fixture recent 2026-07-26T12:00:00Z → also before cutoff → pruned (bug: fixture aged out)
Result PruneOldNeighborMetrics returns 2, test asserts 1 → FAIL: expected 1 row pruned, got 2

Reproduced twice on the unmodified baseline before any change (both runs: observer_neighbor_metrics_test.go:128: expected 1 row pruned, got 2).

This is baseline test nondeterminism, not a production defect — the retention query itself is correct and is unchanged by this PR.

Fix

The production change is exactly one thing: a clock-hook whose default is time.Now. PruneOldNeighborMetrics's retention SQL (DELETE FROM observer_neighbor_metrics WHERE timestamp < ?), its retentionDays parameter, and the exact-boundary (<, not <=) semantics are byte-for-byte unchanged. The only thing that changed is where "now" is read from:

var pruneNeighborMetricsNow = time.Now   // new; production always leaves this alone
...
cutoff := pruneNeighborMetricsNow().UTC().AddDate(0, 0, -retentionDays).Format(time.RFC3339)  // was: time.Now()

Smallest deterministic correction, following the order of preference given for this task:

  • No existing injected clock was available for this function (checked PruneOldMetrics, PruneDroppedPackets, PruneOldClientReceptions — all call time.Now() directly, no seam).
  • Sibling tests (TestPruneOldMetrics, TestPruneOldClientReceptions) already avoid hardcoded dates by computing fixtures relative to time.Now() at test-run time. That pattern is safe for "clearly older" / "clearly newer" cases, but not safe for pinning the exact boundary: the test's own time.Now() call and the production function's internal time.Now() call happen a few instructions apart, and since storage truncates to RFC3339 (second) precision, an unlucky second-boundary straddle could occasionally flip the boundary row's fate — a real, if rare, source of flakiness that would violate "deterministic regardless of execution speed."
  • This package already has an established idiom for exactly this problem: readProcSelfIOFn in stats_file.go, a package-level function variable swapped in tests and restored via t.Cleanup. This PR adds one narrowly-scoped clock hook, pruneNeighborMetricsNow, following that identical pattern, used only by PruneOldNeighborMetrics.

No sleeps, no broadened tolerances, no retention-period changes, no date-relabeling to another future date, and no test skip/quarantine.

Exact-boundary semantics

PruneOldNeighborMetrics's query is WHERE timestamp < cutoff — a strict less-than. So a row exactly retentionDays old (timestamp == cutoff) is retained, not pruned. The rewritten test pins this explicitly:

fixture age relative to fixed now expected
pkOld retentionDays + 1 days pruned
pkBoundary exactly retentionDays retained (pins the < vs <= semantic)
pkRecent retentionDays − 1 days retained
pkRecentOffset same instant as pkRecent, expressed with a +02:00 (Copenhagen) offset instead of Z retained, identically to pkRecent (timezone-invariance: normalizeReportTS re-stores everything as canonical UTC)

Each fixture uses a distinct pubkey so the assertions prove the four rows don't influence one another's fate (not just an aggregate count).

Clock-hook safety review

A second review pass specifically re-examined whether the new package-level pruneNeighborMetricsNow variable could be read by a background goroutine while a test is overriding or restoring it — absence of t.Parallel() alone is not sufficient evidence, so this traces full ownership/lifecycle, not just that one fact:

  • Every reader, enumerated: pruneNeighborMetricsNow is read at exactly one call site (db.go, inside PruneOldNeighborMetrics). PruneOldNeighborMetrics itself is called from exactly three places in the whole repo: two in main.go (both inside func main()'s startup/ticker code, never reachable from go test), and one inside TestPruneOldNeighborMetrics itself.
  • No async migration ever reaches it: OpenStore/OpenStoreWithInterval schedule two background migrations via RunAsyncMigration (obs_observer_ts_idx_v1, an index build; tx_last_seen_backfill_v1 → backfillTxLastSeen). Read both function bodies: neither calls PruneOldNeighborMetrics or touches pruneNeighborMetricsNow, directly or indirectly. WriterExec/instrumentedExec (the DB call inside PruneOldNeighborMetrics) are fully synchronous — no goroutine is spawned inside the function that could outlive the call and read the hook after t.Cleanup has already restored it.
  • Restoration is guaranteed even on test failure — verified empirically, not just from stdlib docs. Two temporary (uncommitted, deleted immediately after use, never part of this diff) diagnostic tests were run in the same package to check this directly:
    • A test that overrides the hook and then calls t.Fatalf before reaching its normal end: the next test observed the hook correctly restored to time.Now.
    • A test that overrides the hook and then panic()s outright: a side-effecting print placed inside the t.Cleanup callback fired (proving the cleanup ran) before the test binary's crash from the unrecovered panic — and since an unrecovered panic terminates the whole test process, there is no subsequent test execution window in which a "half-restored" hook could ever be observed.
  • Conclusion: no concrete new race risk found. This is based on full ownership tracing (above) and empirical -race evidence: TestPruneOldNeighborMetrics alone, -race -count=20 → 0 races; full cmd/ingestor package -race (both by this work and independently by a separate reviewer) → 0 race reports ever implicate pruneNeighborMetricsNow, PruneOldNeighborMetrics, or TestPruneOldNeighborMetrics in any run. The races that do exist in the package (below) are structurally unrelated — none of them touch this variable.

Two pre-existing, unrelated baseline problems (NOT fixed here — documented, not touched)

These are kept visible, not silenced. This PR does not claim a fully green test suite.

A. TestBackfillTxLastSeen_ResolvesFromMaxObservationTimestamp — timing/assertion flake (no -race needed to see it)

  • Repro: cd cmd/ingestor && go test -run '^TestBackfillTxLastSeen_ResolvesFromMaxObservationTimestamp$' -v -count=10 . — toolchain: local go1.26.0 darwin/arm64 (module declares go 1.22; behavior is unaffected by Go version — this same test also fails on Go 1.22.12 and 1.27.1 per earlier CI-toolchain research).
  • On unmodified fda24ca5: 10 isolated single-process runs → 2 FAIL / 8 PASS; a separate 5-run sample → 3 FAIL / 2 PASS. Flake rate is not fixed — it depends on goroutine scheduling.
  • Exact failure: tx_last_seen_backfill_test.go:63: last_seen = 100, want 300 (MAX of 100,300,200).
  • What is established as fact (captured directly in the log, not inferred): newTestStore/openNeighborsStore call OpenStore, which unconditionally schedules backfillTxLastSeen as an async migration (tx_last_seen_backfill_v1) the instant the store opens. The test also calls backfillTxLastSeen(context.Background(), store.db) directly and synchronously right after seeding. A captured failing run shows the tell-tale interleaving directly in the log — two "Backfilling transmissions.last_seen..." start lines print before either completion line appears, proving two literally-concurrent executions of the same function against the same DB handle:
    [async-migration] "tx_last_seen_backfill_v1" starting (boot continues)
    [migration/async] Backfilling transmissions.last_seen from MAX(observations.timestamp)...   <- call #1 start
    [migration/async] Backfilling transmissions.last_seen from MAX(observations.timestamp)...   <- call #2 start (before #1 has completed)
    [migration/async] transmissions.last_seen backfill complete: 0 rows updated
        tx_last_seen_backfill_test.go:63: last_seen = 100, want 300 (MAX of 100,300,200)
    [migration/async] transmissions.last_seen backfill complete: 1 rows updated
    
    backfillTxLastSeen's own selection filter is WHERE last_seen = 0 ... (db.go), so whichever of the two concurrent calls' UPDATE runs first "claims" the row — if that happens to land in the brief window between the test's INSERT of observation 100 and its later INSERTs of 300/200, the row gets stamped last_seen = 100 and is permanently excluded (last_seen != 0) from ever being recomputed by the test's own later, deliberate call. 10/10 runs (not just the failing ones) show exactly two start/complete pairs of this migration — confirming the double-invocation itself is deterministic; only the exact interleaving/outcome is not.
  • What remains a hypothesis (not proven with a goroutine-ID-level trace): the precise sub-statement timing of which specific INSERT inside the test's seeding loop the async call's SELECT/UPDATE lands between. The mechanism (two unsynchronized concurrent callers of the same "claim-once" function) is established fact; the microsecond-level interleaving on any one run is not individually traced.
  • This PR does not fix this (forbidden scope: race fixes). Proposed follow-up: file a new, separate issue — searched dborup/CoreScope issues (gh issue list --state all, keyword and full-list sweep) and found none covering this; only 4 issues exist in this fork total, none related. Not created without approval.

B. Data race between the async backfill migration and a test's own log-buffer capture (issue1865_test.go) — needs -race, self-contained to one test

  • Repro: cd cmd/ingestor && go test -run '^TestHandleNeighborsReportInvalidTimestampLogsEvenWithoutScopeEvidence$' -race -v -count=50 . — same toolchain as above.
  • On unmodified fda24ca5, running this ONE test completely alone (nothing else in the test binary, ruling out any cross-test leak): 5/50 runs hit WARNING: DATA RACE.
  • Exact race (full stack trace captured):
    WARNING: DATA RACE
    Write at 0x00c000524450 by goroutine 26:
      bytes.(*Buffer).Write()
      log.(*Logger).output() / log.Println()
      github.com/corescope/ingestor.backfillTxLastSeen()                    db.go:246
      github.com/corescope/ingestor.(*Store).RunAsyncMigration.func1()      async_migration.go:125
    
    Previous read at 0x00c000524450 by goroutine 23:
      bytes.(*Buffer).String()
      github.com/corescope/ingestor.TestHandleNeighborsReportInvalidTimestampLogsEvenWithoutScopeEvidence()   issue1865_test.go:731
    
    Goroutine 26 created at:
      ...RunAsyncMigration() -> OpenStoreWithInterval() -> OpenStore() -> openNeighborsStore()   issue1865_test.go:114
      -> TestHandleNeighborsReportInvalidTimestampLogsEvenWithoutScopeEvidence()                 issue1865_test.go:719
    Goroutine 23 created at: testing.(*T).Run() [this is the test's own main goroutine]
    
  • What is established as fact: both racing goroutines' stacks trace back to the same single test function — openNeighborsStore(t) at that test's own line 719 schedules the async backfillTxLastSeen migration, and a few lines later (line 731) that same test does log.SetOutput(&buf) + reads buf.String() without first waiting for its own just-spawned migration goroutine to finish. This is self-contained to one test, not a cross-test leak — confirmed by reproducing it with this test running completely alone, in isolation, with no other test in the binary at all. (This corrects a looser earlier characterization — of a "leaked goroutine colliding with a later test's log capture" — that was proposed before this isolation run; the isolation run shows no other test's involvement is needed or occurs.)
  • What is hypothesis, not fully proven: whether every one of the 8 other tests in this package using the same log.SetOutput-capture pattern (decode_error_log_test.go, default_scope_bench_test.go, ingest_buffer_test.go, mqtt_reconnect_test.go, mqtt_watchdog_r2_test.go, mqtt_watchdog_1810_test.go, multibyte_persist_helpers_test.go) is equally exposed — each was located by grep but not individually stress-tested for this PR; flagged as a strong "likely, not yet each individually confirmed" candidate for the same class of bug in the proposed follow-up below.
  • This PR does not fix this (forbidden scope: race fixes). No existing issue found (same search as above). Proposed follow-up: separate issue for this specific self-contained race, distinct from A. Not created without approval.

Correction re: a third, distinct pre-existing race (found by this review, previously mischaracterized)

While re-verifying B, a full-cmd/ingestor-package -race run captured during this PR's own earlier validation work was re-examined stack-trace-by-stack-trace, and turned out to be a third, separate bug, not an instance of B as originally assumed:

Read at ... by goroutine (StartStatsFileWriter.func1, stats_file.go:263)
Previous write at ... by goroutine (TestStatsFileWriter_SampledAtMatchesProcIOSampledAt's own t.Cleanup, stats_file_timestamp_test.go:52)

This is the same class of bug as B (a leaked async goroutine racing a test's own restore of shared state), but a different variable, a different file pair, and — like B — self-contained to one test: reproduced by running TestStatsFileWriter_SampledAtMatchesProcIOSampledAt alone, -race -count=30, on unmodified fda24ca5 → 7/30 race warnings, 4/30 FAIL. Go's race detector attributes a race report to whichever test happens to be executing on the CPU when it fires, which is why an earlier full-package run's race was attributed to TestBackfillTxLastSeen_ResolvesFromMaxObservationTimestamp (problem A's test, which was mid-run at that moment) even though neither of that race's two actual goroutines belongs to that test at all. There are at least three distinct pre-existing bugs in this package's async-migration test infrastructure (A, B, and this one), not one "flake" or two. This one is also not fixed here; also not yet filed as an issue (no duplicate found; not created without approval).

Scope audit

  • Changed files (exactly 2): cmd/ingestor/db.go (+9/−1: adds the pruneNeighborMetricsNow var, changes one call site from time.Now() to pruneNeighborMetricsNow()), cmd/ingestor/observer_neighbor_metrics_test.go (+77/−14: rewrites TestPruneOldNeighborMetrics).
  • No Go version changed anywhere (this branch's go.mod/go.sum are byte-identical to origin/master).
  • No Dockerfile, Dockerfile.go, or deploy.yml changed.
  • No database schema or migration changed.
  • No other test edited.
  • git diff --check: clean (no whitespace errors) — reconfirmed against exact head 1c6a8afc.
  • gofmt -l on both changed files: clean — reconfirmed against exact head 1c6a8afc.
  • PR chore(deps): bump to golang:1.27.1-alpine3.24 and alpine:3.24 #22's branch, head SHA, and files: unchanged (re-verified again at the end of this update).

Repeated-test results

All runs against this branch's exact head (1c6a8afc), cmd/ingestor, local Go 1.26.0 darwin/arm64 (go.mod declares go 1.22, unaffected by this change):

run result
TestPruneOldNeighborMetrics alone, -count=25 (+ a fresh reconfirm -count=10) 35/35 PASS
TestPruneOldNeighborMetrics alone, -race -count=20 (+ a fresh reconfirm -count=10) 30/30 PASS, 0 races
TestPruneOldNeighborMetrics alone, TZ=UTC, -count=5 5/5 PASS
TestPruneOldNeighborMetrics alone, TZ=Europe/Copenhagen, -count=5 5/5 PASS
Negative control (boundary fixture moved 1s past cutoff) FAILS as expected — reverted before commit
Related neighbor-metrics tests (TestRecordObserverNeighborMetrics_*, TestHandleNeighborsReport_RecordsSnrHistory) all PASS
Sibling prune tests (TestPruneOldMetrics, TestPruneOldClientReceptions, TestPruneDroppedPackets) all PASS (unmodified, unaffected)
Full cmd/ingestor suite, no race, -count=1 711 pass, 1 skip, 1 unrelated fail (Problem A)
Full cmd/ingestor suite, -race -count=1 711 pass, 1 skip, 1 unrelated fail + 1 unrelated race (Problem A's test + the third race above, mid-flight). TestPruneOldNeighborMetrics itself: PASS, 0 races, in every run including this one.
cmd/server full suite, -count=1 PASS (28.7s), unaffected
cmd/decrypt, cmd/migrate, all 10 internal/* modules all PASS / [no test files] where applicable
go build ./... + go vet ./..., all 15 Go modules clean

This PR does not claim a fully green suite. Problems A and B (plus the third race documented above) are pre-existing on unmodified fda24ca5, independent of this change, and are explicitly out of scope here (no race fixes, per this task's boundaries) — see the dedicated section above for reproduction commands, exact stack traces, and what's established fact vs. hypothesis for each.

Confirmation

  • Go version: unchanged (no go.mod/go.sum/Dockerfile/deploy.yml diff anywhere in this branch).
  • Docker / deploy / production behavior: unchanged. The only production change is the pruneNeighborMetricsNow clock-hook indirection described above; it defaults to time.Now and is never touched outside this one test.
  • This PR does not approve, endorse, or merge chore(deps): bump to golang:1.27.1-alpine3.24 and alpine:3.24 #22.
  • No merge or deployment was performed as part of this work.

🤖 Generated with Claude Code

Actions / CI status (observed only — nothing restarted or approved)

  • This PR's own CI run (run 34750125325): cmd/ingestor is fully green — ok github.com/corescope/ingestor 116.339s coverage: 76.7% of statements, no FAIL line for any Go test, and TestBackfillTxLastSeen_ResolvesFromMaxObservationTimestamp (Problem A) did not trigger in this particular run, consistent with it being a probabilistic flake rather than a deterministic failure. The only failing job on this PR is an unrelated frontend JS test — ✗ Exactly one 'api('/scope-stats'' call exists (the fixed loader) — found 2 (#1375 FAIL) — in a file untouched by this branch; present identically on unmodified origin/master since this branch never touches any JS file.
  • Independent, live corroboration that Problem A is actively breaking CI today, right now, for unrelated work: several other open PRs on this same fork (codex/port-upstream-1969-..., codex/port-upstream-1957-..., codex/port-upstream-1962-..., etc. — unrelated branches, touching none of the files in this PR) show --- FAIL: TestPruneOldNeighborMetrics and/or --- FAIL: TestBackfillTxLastSeen_ResolvesFromMaxObservationTimestamp in their own CI runs from the last few hours (gh run list, gh run view --log-failed). This is real-world confirmation, independent of this PR, that Problem-A-the-hardcoded-fixture-date is now actively red-flagging CI for anyone building on master, not just a theoretical/local finding.
  • Marking this PR "ready for review" cannot trigger a new workflow run, let alone a deploy — checked directly against the live deploy.yml: on: pull_request: branches: [master] has no types: key, so per GitHub's default it only listens for opened, synchronize, reopened — not ready_for_review. Independently, the deploy job is separately gated if: (github.event_name == 'push' || github.event_name == 'workflow_dispatch') && github.ref == 'refs/heads/master', so it cannot run for a pull_request event or a non-master branch regardless. No workflow was restarted or approved as part of this update.

Review verification (independent pass on exact head 1c6a8afc)

Base fda24ca5 (unchanged origin/master), isolated checkout, local go1.27.0 darwin/arm64, all runs sequential. The line counts in the scope audit above were corrected to the actual diff: db.go +9/−1 and the test +77/−14. No code or head change was made.

Clock hook: no concrete new risk found

  • Readers and writers. pruneNeighborMetricsNow is read at exactly one place (db.go:2355, inside PruneOldNeighborMetrics). It is written only in TestPruneOldNeighborMetrics: the override at line 148, and the restore via t.Cleanup at line 147, which is registered before the override.
  • Callers. PruneOldNeighborMetrics is called in production only inside func main(): at startup (main.go:270) and in the 24 h ticker goroutine (main.go:338). No test calls main(), and no other test calls the function, so no other goroutine can read the hook during the test, independent of t.Parallel().
  • Background goroutines. openNeighborsStore → OpenStore starts two RunAsyncMigration goroutines: the obs_observer_ts_idx_v1 index build and backfillTxLastSeen. Neither reaches the prune function or the hook. Store.Close waits for them via backfillWg.Wait(). Cleanups run in reverse order, so the hook is restored before the store closes. PruneOldNeighborMetrics → instrumentedExec → WriterExec is synchronous.
  • Restore on pass and on failure. t.Cleanup runs on normal completion and on t.Fatal/FailNow. The panic case is covered by the author's empirical check above. A failure before the override only restores the original value.
  • Unchanged behaviour. The default is still time.Now. Within the function, only the cutoff line differs from master. The SQL (DELETE FROM observer_neighbor_metrics WHERE timestamp < ?), the retentionDays parameter, the signature and both main.go call sites are unchanged. db.go is the only production file in the diff.
  • Boundary. RecordObserverNeighborMetrics stores normalizeReportTS(ts) = t.UTC().Format(time.RFC3339), the same fixed-width format as cutoff, so the string comparison is chronological. Rows older than the cutoff are deleted, and a row exactly at the cutoff is kept (strict <). Changing < to <= would fail both the n != 1 and the pkBoundary assertions. The negative control documented above (boundary fixture moved 1 s past the cutoff → fails) was not repeated.
  • Nit, no correctness impact. The offset fixture is labelled "Copenhagen/CEST +02:00", but the fixed instant is in January, when Copenhagen is +01:00. It is a time.FixedZone, so the arithmetic is correct.

Fresh results

Check Result
gofmt -l on both files clean
git diff --check fda24ca5..1c6a8afc clean
cmd/ingestor go build + go vet ./... at head pass
TestPruneOldNeighborMetrics on master, TZ=UTC and TZ=Europe/Copenhagen FAIL in both: expected 1 row pruned, got 2
same test at head, TZ=UTC, -count=20 20/20 pass
same test at head, TZ=Europe/Copenhagen, -count=20 20/20 pass
same test at head, TZ=UTC, -race -count=20 20/20 pass, 0 race reports
same test at head, TZ=Europe/Copenhagen, -race -count=20 20/20 pass, 0 race reports
related tests (TestRecordObserverNeighborMetrics_*, TestHandleNeighborsReport_RecordsSnrHistory, TestPruneOldMetrics, TestPruneOldClientReceptions, TestPruneDroppedPackets) at head 8/8 pass
any race report in the head logs mentioning pruneNeighborMetricsNow / PruneOldNeighborMetrics none

Pre-existing problems this PR does not fix (kept separate)

A. Functional backfill timing failure: TestBackfillTxLastSeen_ResolvesFromMaxObservationTimestamp, no -race needed

  • Fresh runs, -count=20: master failed 1/20, this head 4/20. Every failure had the same assertion: last_seen = 100, want 300 (MAX of 100,300,200).
  • Why the rates differ is not attributed to this PR. It touches neither backfillTxLastSeen nor that test. The author's earlier master samples (2/10 and 3/5) span the same range.
  • Uncertainty. The samples are small, so the rate is not established, and the exact interleaving on a given run is not traced (see A above).
  • CI. It has also failed in CI on several split PRs whose cmd/ingestor tree equals master's.

B. Race-detector data race against the test's log buffer: TestHandleNeighborsReportInvalidTimestampLogsEvenWithoutScopeEvidence, -race

  • Fresh runs, -race -count=50: master had 6 race reports (3 failing iterations), this head 6 race reports (4 failing iterations).
  • Every failure is the race detector's own "race detected during execution of test". The signature is identical on both: log.Println → bytes.(*Buffer).Write from backfillTxLastSeen/RunAsyncMigration, racing bytes.(*Buffer).String at issue1865_test.go:731.
  • No report involves the clock hook.
  • Uncertainty. The race rate is probabilistic. Whether the other log.SetOutput-capturing tests are equally exposed is not individually verified here.

The third race described above (stats_file writer vs stats_file_timestamp_test.go cleanup) was not re-run in this pass.

CI (observed only; nothing restarted or approved)

  • Run 34750125325 on 1c6a8afc: “Go Build & Test” = failure.
    • Steps that passed: Build and test Go server, Build and test Go ingestor, channel/decrypt, Dockerfile invariants, disk-monitor and CSS lint.
    • The job fails at Run JS unit tests: ✗ Exactly one api('/scope-stats' call exists … found 2 (#1375 FAIL). That file is untouched here, and the test also fails on unmodified master locally.
    • Because the step uses set -e, later JS tests and the remaining steps in that job did not run.
    • Playwright E2E, image build, release artifacts, badges and Deploy Staging were all skipped.
  • The CI result is therefore not green, and this PR does not claim a green suite.
  • Ready status. The PR was already marked ready for review before this pass, so no state change was made. pull_request has no types: key, so ready_for_review and body edits do not start runs, and the deploy job requires push/workflow_dispatch on refs/heads/master.

Not run in this pass

  • A full fresh cmd/ingestor/cmd/server suite (only the targeted runs above).
  • The third race above.
  • Playwright E2E and eslint.

TestPruneOldNeighborMetrics hardcoded its "recent" fixture to
2026-07-26T12:00:00Z. Once the real date passed 2026-08-25 (30 days
later), PruneOldNeighborMetrics(30) started pruning that row too,
failing the test's "expected 1 row pruned, got 2" assertion on every
run since -- reproduced twice against unmodified origin/master
(fda24ca) before this change.

This is baseline test nondeterminism, not a production defect: the
retention query itself (`WHERE timestamp < cutoff`) is unaffected and
unchanged.

Fix: pin the test to a fixed reference instant via a new
pruneNeighborMetricsNow package-level clock hook (same swap-in-test /
restore-in-cleanup pattern already used by readProcSelfIOFn in
stats_file.go), instead of a hardcoded calendar date. Production keeps
calling time.Now() by default -- only the indirection changes, not the
retention behavior, the 30-day window, or any caller.

The rewritten test also pins the exact retention boundary explicitly
(a row exactly `retentionDays` old is retained, matching the query's
strict `<`), proves multiple rows don't influence each other's fate,
and proves the fixture's timestamp timezone offset doesn't change the
outcome (normalizeReportTS re-stores everything as canonical UTC).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This was referenced Sep 13, 2026
@dborup
dborup marked this pull request as ready for review September 13, 2026 10:47
@adminopenclaw8-sketch
adminopenclaw8-sketch merged commit 6b70e94 into master Sep 13, 2026
5 of 6 checks passed
adminopenclaw8-sketch pushed a commit that referenced this pull request Sep 14, 2026
…rl comments

Two review fixes on this PR.

1. CI registration. `test-issue-1890-og-url.js` was only wired into
   `test-all.sh` (used by `npm test`), not into the JS test step in
   `.github/workflows/deploy.yml`. CI could not catch a regression of
   the hardcoded og:url. Added `node test-issue-1890-og-url.js` to that
   step's existing list, directly before `test-issue-1375-scope-stats-
   fetch.js` (the test that currently stops the step). test-all.sh is
   unchanged; the registration there was already correct.

2. Comment accuracy, in `public/index.html` and
   `test-issue-1890-og-url.js`. The removed tag declared the upstream
   analyzer instance's own URL as the canonical URL for every
   self-hosted deployment (Kpa-clawbot#1890) -- correct metadata for upstream,
   wrong for everyone else. Reworded both comments to say that
   precisely, and to stop implying things not established:
   - og:url is metadata, not an HTTP redirect, and does not by itself
     decide what a viewer's click navigates to.
   - Per the Open Graph protocol it is a required property, not
     "optional" -- omitting it leaves the crawled URL as the fallback
     canonical reference, which is what actually fixes this for every
     instance.
   - This change does not refresh previews a consumer has already
     cached under the old, hardcoded value.
   No functional assertions changed in test-issue-1890-og-url.js --
   only the file-level comment. The `public/index.html` comment change
   had to avoid writing the literal removed domain: the test's own
   4th assertion scans the whole file for that string, and an earlier
   draft of this comment briefly reintroduced it and failed its own
   guard before landing on the current wording.

Verified on this branch's own base (pre-#51/#33 master) and against
a merge into current master: test-issue-1890-og-url.js passes 4/4 on
the resulting index.html and still fails 2/4 (og:url present, 00id.net
present) against the original hardcoded tag. The merge result's
deploy.yml differs from current master by exactly the one added test
line; release-fast-path.yml and cmd/server/fork_guard_workflow_test.go
are byte-identical to master, and all four #51 repository guards
(the five build-and-publish publish steps, release-artifacts, deploy,
publish, retag-or-fallback) are present in the merged file.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants