Repository navigation
feat(ingestor): resolve the last hop from the observer (#188) - #190
Conversation
#188) A non-advert observation has no anchor for its first hop, so with 1-byte hashes (about 7 nodes per prefix on staging) every hop stayed nil and the row was stored as NULL. The observer is a better anchor: a flood forwarder appends its hash before retransmitting (MeshCore Mesh.cpp routeRecvPacket) and the observer logs the packet before its own routing step (Dispatcher.cpp logRx), so the last hash is the node it heard, a direct neighbour. resolveObservationPath keeps the forward chain from fromPubkey and, for flood route types only, adds a backward chain: the last hop is the unique neighbour of the observer, each earlier hop the unique neighbour of the hop after it. The merge keeps every forward hop and fills nil hops from the backward chain; if the two disagree on a hop, or would place one node twice, the backward result is dropped for the row. No geo, GPS or count tie-break. DIRECT and TRACE paths are not anchored. path_resolver.go is gofmt-ed as a whole (one comment table moved). Relates to #188, #184 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The ingestor's prefix index held every node, so a prefix shared by one repeater and a companion counted as ambiguous, and a prefix unique to a companion resolved to it (64 companion pubkeys are stored as relays on staging). buildPrefixIndex now keeps only nodes for which isRelayRole is true: the server's canAppearInPath (repeater, room_server, room), with the same test cases. Companions and sensors ship with forwarding off in the firmware (companion_radio NodePrefs.h, simple_sensor SensorMesh.cpp); a companion that opted in to repeating stays unresolved, as on the server. The index also feeds the neighbour-edge builder, so edges to companions are no longer derived from a hop prefix. Existing fixtures that seeded hop nodes without a role now give them role 'repeater'; expectations are unchanged. Relates to #188, #184 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rimed (#188) Rows stored before the observer anchor and the relay-only index, and rows drained from the start-up buffer before the prefix index is primed, keep resolved_path = NULL. RunResolvedPathBackfill re-resolves them once per ingestor start, with resolveObservationPath, for observation ids between the persisted watermark and MAX(id) at the start of the pass. Each batch reads at most batchSize rows by primary key without writerMu, resolves them outside any lock, and commits one WriterTx that sets resolved_path only where it is still NULL and moves the watermark (resolved_path_backfill_state, one row). A restart repeats at most the uncommitted batch. Batches are paced by a pause. Defaults: 500 rows, 250 ms, configurable as resolvedPathBackfill.{batchSize,pauseMs,disabled}. main.go starts the pass right after StartNeighborEdgesBuilder has primed the index and graph. The server indexes a row when it first polls it or at Load, so backfilled values reach its indexes on its next restart; the server is unchanged. Relates to #188, #184 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… is primed (#188) Started before StartNeighborEdgesBuilder had primed the prefix index and graph, a pass would resolve nothing yet move its watermark past every row, so those rows would never be retried. RunResolvedPathBackfill now returns an error and leaves the watermark unchanged until both are loaded. Relates to #188 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Review — CS-MacBook review PR#190 observer-anker — head 1cfb42dDom: APPROVE med nits Independent, read-only review of head Evidence tags: [F] measured against the staging public API, [T] test or benchmark I ran, [A] analysis from reading code or firmware, [K] not checked. SummaryI found no correctness bug in the resolver, the role filter or the backfill that blocks the merge. The firmware reasoning holds, the merge rule is sound for identity and duplicates, and live ingest is not starved. The deliberate scoring change behaves as described: my own replay on a fresh staging sample reproduces the PR's NULL-share numbers. Three things I would fix or document before merging (findings 1 to 3), plus smaller nits. Findings
Answers to the review points1. Resolver correctness
2. Relay-role filter
3. Backfill
4. Performance
5. Effect (fresh replay)I fetched the nodes list, the neighbour graph and 8 windows of 150 transmissions (about 21,940 path observations over ~36 h) from the staging public API, a handful of GET requests, and ran this branch's Go code over them with the graph proxy.
[F][T] The old-resolver replay reproduces the persisted NULL share (27.6 % vs 28.0 %), so the replay is credible, and the new-resolver numbers agree with the PR's (5.0 % NULL, 13.3 % for 1-byte, 4.0 % disagreement). My ratio is 1.35 instead of 1.47 because my sample is biased towards well-observed transmissions. Absolute levels differ from the PR (33 vs 17.5) for the same reason. The expected Traffic Share sum of about 22-26 is plausible as an end state. 6. Conflict with #141
Tests
Mutants (each in a copy of the PR tree, then the targeted suite)
CI (head
|
…y graph (#188) PR #190 review, finding 1. StartNeighborEdgesBuilder loads the graph snapshot before its warm-up build, and the backfill only checks that the graph is non-nil. With empty neighbor_edges the pass resolves only unique prefixes and moves its watermark past every row for good. Three tests, red by assertion on 1cfb42d: - the pass right after StartNeighborEdgesBuilder uses the pre-build graph; - a pass and a single batch on an empty graph commit a watermark; - the background pass started before the first build runs at once. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QbNdgonR8kq7SPEU7LcmXD
…raph (#188) PR #190 review, finding 1. The backfill could run on the graph snapshot that StartNeighborEdgesBuilder loads before its warm-up build, or on an empty graph. It then resolved only unique prefixes and moved its watermark past every other row for good. - StartNeighborEdgesBuilder publishes the graph again after a successful warm-up build, and after each successful tick, as a post-build snapshot (neighborGraphHolder.storeBuilt). - The pass and every batch run only with a non-empty prefix index and a post-build graph that has at least one edge; otherwise they write nothing, not even the watermark (errResolvedPathBackfillNotReady). - StartResolvedPathBackfill waits for the next post-build graph and retries; the retry resumes at the persisted watermark. - Test setup primes the way the builder does: build, then publish. - TestResolvedPathBackfillReady_188 pins each readiness condition. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QbNdgonR8kq7SPEU7LcmXD
…t rule (#188) PR #190 review, finding 2. The three keys were missing from config.example.json. They now sit in a resolvedPathBackfill block with a _comment, like the other ingestor blocks, and in the configuration guide. pauseMs: 0 cannot switch the pause off, because 0 means the default. That is documented rather than changed: a pass with no pause at all would hold the single write connection back to back. The smallest pause is 1 ms. Tests: - TestConfigExample_ResolvedPathBackfill_188 parses config.example.json and checks the block and its defaults (red before the block was added); - TestResolvedPathBackfillSettings_ZeroMeansDefault_188 pins 0 and negative values to the defaults and 1 to 1 ms. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QbNdgonR8kq7SPEU7LcmXD
…188) PR #190 review, finding 5. Three mutants survived the whole ingestor suite. Each test passes on the code and fails on its mutant: - m9, the backward walk keeps its anchor across a nil hop: TestObserverAnchor_BackwardWalkBreaksOnNilHop_188 (the reviewer's [a1 b2 c3] probe, hop 1 ambiguous). - m10, the backward walk without its exclusion set. It was reported as practically equivalent, but it is not: when a prefix repeats, the exclusion leaves one candidate where there would be two. TestObserverAnchor_BackwardWalkExcludesResolvedHops_188. - m12, the backfill passes from_pubkey for non-ADVERT rows: TestResolvedPathBackfill_FromPubkeyOnlyForAdverts_188. Today only ADVERTs carry from_pubkey, so on real data this mutant has no effect. The test keeps the backfill consistent with InsertTransmission if that changes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QbNdgonR8kq7SPEU7LcmXD
…warm-up (#188) Follow-up to the finding-1 fix. If the warm-up build fails, the backfill waits for the first successful tick. The mutant "a tick never publishes a post-build graph" survived the targeted suite; this test kills it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QbNdgonR8kq7SPEU7LcmXD
…188) PR #190 review, finding 4. In a flood path an observer reports, the last hop is the node it heard, never itself. With the observer's prefix unique, both chains name the observer: path ["0b"] resolves to obs188. With path ["c3","0b"], the backward walk then anchors hop 0 on it. TestObserverAnchor_ObserverIsNeverItsOwnLastHop_188 is red by assertion. TestObserverAnchor_ObserverMayBeAnEarlierHop_188 is green and pins the scope: an observer logs a packet before de-duplication, so it can report an echo of a flood it forwarded, with its own hash earlier in the path. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QbNdgonR8kq7SPEU7LcmXD
…#188) PR #190 review, finding 4. A radio does not receive its own transmission, so the last hop of a flood path an observer reports is never the observer. Both chains now leave that hop nil instead: - the forward chain (resolvePathForward, the walk behind resolvePathWithContext) does not take the observer as a candidate for the last hop of a flood path. With a unique prefix, the forward chain alone named the observer, so a fix in the backward walk alone would not change the stored row; - the backward walk excludes the observer from the last hop only. It may still be an earlier hop: it logs a packet before de-duplication (Dispatcher.cpp logRx before processRecvPacket), so it can report the echo of a flood it forwarded. DIRECT routes and resolvePathWithContext's other callers are unchanged. Perf (hot path, InsertTransmission): BenchmarkResolveObservationPath_188 is new, because the PR's benchmark was not in the repo. It uses 700 relays, 60 observers and 1-5 hop 1-byte flood paths; one op is one observation. Run interleaved, n=12 each, on the same machine: - before: median 3,875 ns/op, 119 B/op, 5 allocs/op; - after: median 3,890 ns/op, 119 B/op, 5 allocs/op. That is within noise. The change adds one map insert per flood observation. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QbNdgonR8kq7SPEU7LcmXD
PR #190 review, finding 8. A TRACE's header path carries SNR bytes, but the ingestor's decoder replaces the hops with the planned route from the payload. TRACE is not anchored because it is always DIRECT: sendFlood refuses it. The comment now says so. Firmware references are checked against MeshCore a366955, the version cloned here; the PR cited 0679dbe: - routeRecvPacket is at Mesh.cpp:344-356, with copyHashTo at :349; - removeSelfFromPath is at :334; - the front-of-path match is at :89; - the SNR append is at :60-61; - logRx is at Dispatcher.cpp:238, before processRecvPacket at :246/:256. Comments only. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QbNdgonR8kq7SPEU7LcmXD
|
Review of head
These are code-path findings; I have not run a staging replay or changed the PR. The Go CI job passed on this head, while the Playwright job was still in progress at my last check. I also noticed that #182 has since landed on master, so the final integration should be verified against the current base. |
Brings in #182 (9d29dae), #191, #194 and #196, so that the observer anchor and the backfill are tested against the current base. Master's server now indexes live observations from the persisted resolved_path (#182). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QbNdgonR8kq7SPEU7LcmXD
Rapport — CS-pve-agent1 PR#190 review-rettelser — head 2e2886aStatus: Findings 1, 2, 3, 4, 5 and 8 are fixed, and findings 6 and 7 are assessed with no change. CI is green on This follows up on the review comment "Review — CS-MacBook review PR#190 observer-anker — head 1cfb42d". I built on Evidence tags:
Findings
1. Backfill on a pre-build or empty graphChange:
Red before, green after [T]. The three reproduction tests from All three pass from
Test-setup change: 2. Config
3. PR bodyWhat master actually does [A] (checked in
Body changes:
4. The observer as its own last hopWhy a backward-only fix is not enough [T]: the suggested Fix:
Red before, green after [T]. On It passes from Perf [T]. The resolver is the
That is within noise; the fix adds one map insert per flood observation. 5. Surviving mutants
6. Seam adjacency: assessment, no change [A]A seam without an edge occurs only when the forward chain found no neighbour of hop i among the candidates of hop i+1, and the backward chain then picked hop i+1 from its own anchor. A seam check would be stricter than either chain's own rule, for three reasons:
A seam check would mostly discard correct backward results on 1-byte meshes. I recommend leaving it as is and, if wanted, counting such rows on staging later. 7.
|
| Reference | Location |
|---|---|
routeRecvPacket |
Mesh.cpp:344-356, with copyHashTo at :349 |
removeSelfFromPath |
:334-342 |
| front-of-path match | :89 |
| SNR append | :60-61 |
sendFlood refuses TRACE |
:638 |
logRx |
Dispatcher.cpp:238, before processRecvPacket at :246/:256 |
These are now in the path_resolver.go comment and the PR body.
The TRACE comment now says what the ingestor actually does. The header path carries SNR bytes, but decoder.go replaces the hops with the planned route from the payload. TRACE is not anchored because it is always DIRECT.
Mutants (this follow-up)
Each mutant was applied to a copy of the tree and run against the affected tests (-run '188|ResolvedPath|NeighborEdges|NeighborGraph|PrefixIndex|Resolve'). All 14 were caught [T]. One more test failed in every copy, TestConfigExample_ResolvedPathBackfill_188: the copy had no config.example.json. That failure is not counted.
| Mutant | Caught by |
|---|---|
| f1a: no post-build refresh after the warm-up | …_UsesGraphFromFirstEdgeBuild_188, …_EmptyGraphKeepsWatermark_188, …_StartWaitsForFirstEdgeBuild_188 (whole suite) |
| f1b: readiness ignores "post-build" | TestResolvedPathBackfillReady_188, …_StartWaitsForFirstEdgeBuild_188 |
| f1c: readiness ignores an empty graph | TestResolvedPathBackfillReady_188, …_EmptyGraphKeepsWatermark_188 |
| f1d: readiness ignores an empty index | TestResolvedPathBackfillReady_188 |
| f1e: no per-batch guard | …_EmptyGraphKeepsWatermark_188 |
| f1f: the background pass does not wait | …_StartWaitsForFirstEdgeBuild_188 |
| f1g: a tick never publishes a post-build graph | TestNeighborEdgesBuilder_TickPublishesBuiltGraphAfterFailedWarmUp_188 (survived until that test was added) |
| f4a: no last-hop exclusion in the forward chain | …_ObserverIsNeverItsOwnLastHop_188 |
| f4b: no last-hop exclusion in the backward walk | …_ObserverIsNeverItsOwnLastHop_188 |
| f4c: the backward walk excludes the observer at every position | …_ObserverMayBeAnEarlierHop_188 |
| f4d: the exclusion is also applied to DIRECT routes | …_ObserverIsNeverItsOwnLastHop_188 |
| m9, m10, m12 | see finding 5 |
Local runs
Run on 2e2886af [T]:
| Command | Result |
|---|---|
cd cmd/ingestor && go test -count=1 ./... |
ok (548 s) |
cd cmd/server && go test -count=1 ./... |
ok (314 s) |
go test -race -count=3 -timeout 90m -run '188|ResolvedPath|NeighborEdges|NeighborGraph|PrefixIndex|Resolve' |
ok (705 s), no race reported |
go vet ./... (ingestor and server) |
clean |
gofmt -l on every changed Go file |
clean |
sh test-all.sh |
all passed |
The first -race -count=3 attempt hit go test's default 10-minute timeout inside the existing TestNeighborEdgesBuilderDeltaScan, with no race and no failure. The re-run with -timeout 90m passed.
CI (head 2e2886a)
Run 37129803920, pull_request event [T]:
| Job | Result | Time |
|---|---|---|
| Go Build & Test | success | 21 min |
| Playwright E2E Tests | success | 19 min |
| Build & Publish Docker Image | success | <1 min |
| Release Artifacts | skipped (PR) | – |
| Deploy Staging | skipped (PR) | – |
| Publish Badges & Summary | skipped (PR) | – |
Not verified
- [K] Staging behaviour after deploy: the sum settling, the backfill's duration and hold time on the real DB. None of these changes alter what the review measured there.
- [A] The tests show that the backfill waits for the post-build graph, but on a mesh where no build ever yields an edge, it waits indefinitely. It logs one "waiting" line, and cancelling it on shutdown is covered by the stop function; there is no separate test of that indefinite wait.
- New observation, not changed [A]. Neither the observer anchor nor the neighbour builder checks
observations.direction, whileclient_reception.go(deriveHeardKey) requiresrx.- If an uploader sends
txobservations of floods its node forwarded, the last hop of those paths is the observer itself. - With finding 4 that hop is never resolved to the observer itself. But if exactly one neighbour of the observer shares the observer's prefix, the backward walk picks that neighbour and anchors earlier hops on it: a mis-attribution.
- I did not check whether any uploader sends
tx. It could be a small follow-up: anchor onlyrx, or rows with no direction.
- If an uploader sends
- [A] Anomalous FLOOD-routed TRACE packets get their hops replaced with the planned route (
decoder.go), and the anchor would treat them as a flood. The firmware never sends one (sendFloodrefuses TRACE). Not changed. - [K] feat(ingestor): resolve the last hop from the observer, relay-only prefix index, backfill NULL resolved_path #188 point 3 (the start-up window in
main.go) stays with the original author.
…h few edges (#188) PR #190, second review, P2-1. The warm-up loop in StartNeighborEdgesBuilder stops once buildAndPersistNeighborEdges returns fewer than neighborBuilderMaxBatch edges, but each call reads up to that many observations. A full batch that yields few edges ends the warm-up with observations left unscanned. The backfill then runs on that partial graph and moves its watermark past rows it could not resolve. Two tests, red by assertion: - one full batch (50,000 rows) with a single edge, and the edge the rows need in the second batch: the warm-up stops early; - a full batch with no edge newer than the watermark, so the build cannot move past it: the backfill runs anyway, because the graph holds an older edge. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QbNdgonR8kq7SPEU7LcmXD
…188) PR #190, second review, P2-1. buildAndPersistNeighborEdges reads at most neighborBuilderMaxBatch observations per call, but the warm-up loop stopped once a call returned fewer edges than that. A full batch with few edges ended the warm-up with observations unscanned, and the backfill then ran on that partial graph. - buildNeighborEdges reports both the edge upserts and the observation rows read; buildAndPersistNeighborEdges keeps its old result for existing callers. - The warm-up continues while a call reads a full batch. It stops when a call reads fewer rows: it has caught up. - A full batch with no edge cannot move the watermark (MAX(last_seen)), and the next call would read the same rows. The warm-up then stops, logs it, and leaves the graph unpublished. - Only a build that caught up publishes a post-build graph, in the warm-up and in a tick. The backfill therefore never runs on a graph the builder has not finished, and waits instead. - Comments, config.example.json and the guide now say "caught up". TestNeighborEdgesBuilder_WarmUpScansPastFullBatchWithFewEdges_190 and TestResolvedPathBackfill_WaitsWhileEdgeBuildCannotCatchUp_190 (red in e369ae6) pass. The second one now also lets ticks run. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QbNdgonR8kq7SPEU7LcmXD
PR #190, second review, P2-2. The neighbour builder adds an observer <-> last-hop edge for every observation, whatever its route type. For a DIRECT packet the path is the remaining planned route: forwarders strip themselves from the front (Mesh.cpp:89, removeSelfFromPath :334-342), and only flood forwarders append their hash (routeRecvPacket, :346-350). The last hop of a DIRECT path is therefore the route's far end, not the node the observer heard. TestNeighborEdgesBuilder_ObserverEdgeOnlyForFloodRoutes_190 mixes flood, transport flood, DIRECT and TRANSPORT_DIRECT observations. It is red by assertion: there are edges from DIRECT paths, and the flood resolver leaves an ambiguous last hop nil because of the false edge. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QbNdgonR8kq7SPEU7LcmXD
…#188) PR #190, second review, P2-2. On a DIRECT route the path is the remaining planned route: - each forwarder matches itself at the front and strips itself (MeshCore a366955, Mesh.cpp:89, removeSelfFromPath :334-342); - only flood forwarders append their hash (routeRecvPacket :346-350). So the last hop of a DIRECT path is the far end of the route, not the node the observer heard. The edge it produced was false, and the flood resolver relies on those edges to pick a unique candidate. The builder now reads transmissions.route_type and adds the observer edge only for TRANSPORT_FLOOD (0) and FLOOD (1). An unknown route type (NULL) is skipped as well. Other edges are unchanged: the originator<->first-hop edge of ADVERTs, and the interior hop-to-hop edges. Consecutive hops of a DIRECT route are neighbours. This changes neighbor_edges for every consumer: the Neighbors panel, bridge nodes and the estimates built on them. The server reads the table and has no other production edge source; its extractEdgesFromObs has the same rule but no caller outside tests. DIRECT edges already persisted stay until the neighbour-edge prune drops them (5 days by default). TestNeighborEdgesBuilder_ObserverEdgeOnlyForFloodRoutes_190 (red in the previous commit) passes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QbNdgonR8kq7SPEU7LcmXD
In the P2-2 test the TRANSPORT_FLOOD row used the same path as the FLOOD row, so a builder that dropped route type 0 still passed. The row now names its own relay (a1a), and the test expects both observer edges. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QbNdgonR8kq7SPEU7LcmXD
Rapport — CS-pve-agent1 PR#190 runde 2b — head 0f4c167Status: P2-1 and P2-2 are fixed, with tests that were red before and green after; all 8 mutants are caught. Master (with #182) is merged in. CI is green on This round addresses the two P2 findings in the review of Push: I built on Merge with master: I first merged Evidence tags:
Findings
P2-1: warm-up completionChange:
Red before, green after [T]. Both tests use a real 50,000-row batch, since Both pass from
Mutants [T]. Each was applied to a copy of the tree and run against the affected tests:
P2-2: observer edges only for flood routesFirmware [A], MeshCore
A DIRECT path is therefore the remaining planned route, and its last entry is the far end of the route, not the node the observer heard. Change: the builder now selects
Red before, green after [T]. The last line is the review's concern made concrete: the false DIRECT edge to
Mutants [T]:
Effect beyond this PR [A]. The PR body now states it:
PR body
Local runs (
|
| Command | Result |
|---|---|
cd cmd/ingestor && go test -count=1 ./... |
ok (541 s) |
cd cmd/server && go test -count=1 ./... |
ok (348 s) |
go test -race -count=3 -timeout 90m -run '_190$|188|ResolvedPath|NeighborEdges|NeighborGraph|PrefixIndex|Resolve' |
ok (1,211 s), no race reported |
go vet ./... (ingestor and server) |
clean |
gofmt -l on the Go files changed in this round |
clean |
sh test-all.sh |
201 passed, 0 failed |
On the whole merged diff, gofmt -l flags only cmd/ingestor/source_status.go, which comes from master unchanged. No new interface{} was added in cmd/ingestor, and cmd/server is untouched by this PR.
CI (head 0f4c167)
Run 37135018260, pull_request event [T]:
| Job | Result | Time |
|---|---|---|
| Go Build & Test | success | 22 min |
| Playwright E2E Tests | success | 22 min |
| Build & Publish Docker Image | success | 1 min |
| Release Artifacts | skipped (PR) | – |
| Deploy Staging | skipped (PR) | – |
| Publish Badges & Summary | skipped (PR) | – |
Not verified
- [A] A builder that cannot catch up waits indefinitely. Its warm-up logs one line; ticks stay stuck silently, which is how the builder behaved before. The backfill then never runs. That is deliberate: no watermark moves on a partial graph. A builder whose 50,000-row batch yields no edge would need its own fix, for example a persisted scan watermark; that is outside this PR.
- [A] More full batches without edges. Excluding DIRECT paths from observer edges means a backlog dominated by DIRECT traffic yields fewer edges per batch, so that case becomes slightly more likely. I have not measured its frequency on staging.
- Staging (left to the lead or the developer): the new
neighbor_edgescontents, the Neighbors panel and bridge counts after the DIRECT edges age out, and the backfill's run time. - [A] Carried over from round 1, unchanged:
observations.directionis not checked by the anchor or the builder;- an anomalous FLOOD-routed TRACE is treated as a flood;
- feat(ingestor): resolve the last hop from the observer, relay-only prefix index, backfill NULL resolved_path #188 point 3 stays with the original author.
Review — CS-cloud PR#190 runde 2b — head 0f4c167Dom: APPROVE med nits Independent, read-only review of head
Evidence tags: [T] test, benchmark or CI; [A] analysis of code or firmware; [K] taken from earlier comments and not run again. FindingsNo blocking finding. All priorities are low.
1. Earlier findings: are they really fixed?
2. Effect of P2-2 on other
|
| Run | Result |
|---|---|
cd cmd/ingestor && go test -count=1 ./... |
ok (139 s) [T] |
go test -race -count=3 -run '_190$|_188$|ResolvedPath|ObserverAnchor|PrefixIndex' . (ingestor) |
ok (691 s), no race [T] |
cd cmd/server && go test -count=1 ./... |
first run: 1 failure, TestHandleNodePaths_PrefixCollision_1352 with 503 "index loading"; second run ok (60.5 s) [T] |
sh test-all.sh |
201 passed, 0 failed (rc 0) [T] |
| CI on the head | Go Build & Test, Playwright E2E and Docker build: success; release, badges and staging deploy: skipped (PR) [T] |
The _1352 failure is the known pre-existing flake, not this PR: cmd/server is identical to master. Reproduced with the same counts on both:
-run '_1352$' (4 tests) |
master 6f7a5e7c |
head |
|---|---|---|
default GOMAXPROCS, -count=30 |
0 failures | 0 failures |
-cpu 1, -count=10 |
39 failures | 40 failures |
[T]
Own mutants (in addition to the authors'), each in a copy of the head and run with -run '_188$|_190$|ResolvedPath|ObserverAnchor|PrefixIndex|ResolveHop|NeighborGraph|ObserverEdge|TestNeighborEdgesBuilder_':
| Mutant | Result |
|---|---|
| X1 tick publishes a post-build graph even when the build errored | survives (finding 1) |
X2 observer id not lower-cased in resolveObservationPath |
killed by TestObserverAnchor_ObserverIsNeverItsOwnLastHop_188 |
| X3 NULL route type gets the observer edge | survives (finding 2) |
X4 backward chain does not exclude fromPubkey |
survives (finding 3) |
| X5 backfill ignores the row's route type (treats every row as DIRECT) | killed by 10 tests, including …_ResolvesRowsIngestedBeforePriming_188, …_UsesGraphFromFirstEdgeBuild_188, …_WarmUpScansPastFullBatchWithFewEdges_190 |
[T]
6. Performance and operations
-
Resolver cost [T]: I added a forward-only twin of
BenchmarkResolveObservationPath_188(same fixture: 700 relays, 60 observers, 1–5-hop 1-byte flood paths) in a scratch copy and ran both interleaved, n=6, on a linux/amd64 container (Xeon 2.1 GHz, 4 vCPU, load average under 1):- forward only: median 623 ns [612–657], 31 B/op, 1 alloc;
resolveObservationPath: median 5,640 ns [5,358–5,848], 119 B/op, 5 allocs.
That is about +5 µs per flood observation on this worst-case fixture (every row is 1-byte and leaves a nil hop, so every row runs the backward walk). The PR's 757 → 2,345 ns is on a staging sample where many rows do not take the backward path. Same order of magnitude: microseconds against a millisecond insert. At staging's ~4–5 observations/s that is well under 0.1 ms of CPU per second.
-
Backfill vs. live ingest [A]: each batch reads by primary key without
writerMu, resolves without a lock, and then holds the write transaction only for at most 500 single-rowUPDATEs plus the watermark. That is milliseconds per 250 ms pause. MacBook's end-to-end measurement (liveInsertTransmissionp99 1.2 ms during a default backfill) [K] fits that. I did not re-measure. -
First start on a 5 GB DB [A]:
- Pass size: the pass covers
(watermark, MAX(observations.id)], so on first start that is every stored observation, not only the NULL ones. At ≤2,000 rows/s, 10–30 M rows take about 1.5–4 h. It resumes from the persisted watermark after a restart. - Prerequisite: it starts only after the neighbour-edge warm-up has caught up. With a persisted
neighbor_edges(staging) that is quick. - Risk: if
neighbor_edgesis empty or far behind (for example a restored copy), the warm-up scans from its old watermark in 50,000-row batches. If one full batch yields no edge, the backfill waits indefinitely, by design. The PR documents this; it logs "initial build cannot move past its watermark". - Deploy note: check for the
[resolved_path_backfill] startinganddonelines. Plan a server restart afterdone, then read the traffic-share sum; a rising sum before that is the transition, not drift.
- Pass size: the pass covers
7. Rules
| Rule | Result |
|---|---|
All writes in cmd/ingestor; cmd/server read-only |
yes: cmd/server and internal/ unchanged; no new INSERT/UPDATE/DELETE outside cmd/ingestor [T] |
No new map[string]interface{} |
0 added lines [T] |
config.example.json has the new keys |
yes [T] |
| Fork guards | 9 × github.repository == 'Kpa-clawbot/CoreScope' on the head and on master [T] |
| No closing keywords in title or body | none found (title: "feat(ingestor): resolve the last hop from the observer (#188)") [T] |
Not verified
- [K] Staging after deploy: the sum settling near the estimate, the backfill's real duration and lock-hold time on the staging host, and the Neighbors/bridge/GPS-sanity numbers after the DIRECT edges age out.
- [K] The authors' mutant tables and the replay numbers; I did not rerun them.
- [A] The real number of observation rows in the 5 GB DB; the duration above is an estimate.
- Not run: the full
cmd/serversuite under-race, by request.
Generated by Claude Code
Relates to #188, #184, #182
This is a deliberate scoring change. On staging, every non-advert observation with 1-byte hop hashes is stored with
resolved_path = NULL, so about half of the relayed traffic gives no relay credit wherever the server indexes from that column. This PR resolves most of those rows. Traffic share and usefulness change for most repeaters. The expected sum oftraffic_share_scoreis about 1.47× today's persisted level (see the replay below).Dependency on #182: now met. #182 is merged on master (
9d29daec), and master is merged into this branch (486aa9ea).indexResolvedPathHops, which feedsbyPathHopand so traffic share) read the persistedresolved_pathonly at start-up load (Load,loadChunk). Live-polled observations used the server's own resolver.indexObservationRelayHopsfeeds those indexes from the persisted column on start-up load and on live polling alike. This PR's resolution therefore reaches the score continuously.What changes (cmd/ingestor only)
1. Observer anchor (
path_resolver.go,db.go)resolveObservationPathreplaces the directresolvePathWithContextcall inInsertTransmission.fromPubkey(ADVERT only) and each next hop on the previous resolved hop.resolveHopWithContext, so a hop resolves only when exactly one candidate qualifies. There is no geo, GPS or count tie-break (fix(nodes): rebuild relay-hop history on startup from path_json Kpa-clawbot/CoreScope#1643).<observer_id>segment of the MQTT topicmeshcore/<iata>/<observer_id>/packets, which is the observer node's pubkey.neighbor_builder.go).NeighborGraphlower-cases on lookup.companion) has no edges and anchors nothing.Firmware basis (MeshCore
src/; line numbers verified at a366955):Identity.h:19-31).Mesh.cpp:344-356,routeRecvPacket,copyHashToat:349).Dispatcher.cpp:238logRx, beforeprocessRecvPacketat:246/:256).Not anchored:
Mesh.cpp:89,removeSelfFromPath:334-342).sendFloodrefuses it (Mesh.cpp:638). Its header path carries SNR bytes (Mesh.cpp:60-61), and the ingestor's decoder replaces the hops with the planned route from the payload.Packet.h:14-17,64-65.2. Relay-role prefix index (
neighbor_builder.go)buildPrefixIndexkeeps only nodes for whichisRelayRoleis true. That is the server'scanAppearInPath(cmd/server/store.go): a role containingrepeaterorroom_server, or equal toroom. It is copied because the two binaries share no package, and its tests mirror the server'sTestCanAppearInPath.On staging, 64 companion pubkeys are stored as relays today.
Firmware: companions and sensors ship with forwarding off:
companion_radio/NodePrefs.h:91disable_fwd = 1, andMyMesh.cpp:891;simple_sensor/SensorMesh.cpp:728.Forwarding is opt-in (
companion_radio/MyMesh.cpp:1402). A hop through a companion that opted in stays unresolved, or resolves to the one relay that shares its prefix. That is the same trade-off the server makes.The index also feeds the neighbour-edge builder, so edges to companions are no longer derived from hop prefixes.
2b. Observer↔last-hop edges from flood routes only (
neighbor_builder.go)The builder used to add an observer↔last-hop edge for every observation. For a DIRECT route the path is the remaining planned route:
Mesh.cpp:89,removeSelfFromPath:334-342);routeRecvPacket:346-350).So a DIRECT path's last hop is the far end of the route, not the node the observer heard, and the edge was false. Since the observer anchor relies on these edges, the builder now adds the observer edge only for route types 0 and 1. A NULL route type is skipped too. The ADVERT originator↔first-hop edge and the interior hop-to-hop edges are unchanged; consecutive hops of a DIRECT route are neighbours.
This changes
neighbor_edgesfor every consumer: the Neighbors panel, bridge nodes and the estimates built on the graph. The server reads the table; itsextractEdgesFromObshas the same rule but no caller outside tests. DIRECT-derived edges already persisted stay until the neighbour-edge prune removes them (5 days by default), so the old edges fade out over that window after deploy.3. Start-up window (
main.go~318 / ~487): not in this PRDeferred until #141 was merged, because #141 rewrites that part of
main.go. #141 (41675d98) has since been merged; point 3 stays with the original author. Meanwhile, the backfill below resolves rows ingested in that window.git merge-treeof this branch with41675d98merges cleanly.4. Backfill (
resolved_path_backfill.go,config.go,main.go)There is one pass per ingestor start. It is started after
StartNeighborEdgesBuilder, and it runs only once a neighbour graph from a successful edge build is published (see the readiness guard). It covers observation ids from the persisted watermark up toMAX(id)at the start of the pass.Each batch:
batchSizerows by primary key, withoutwriterMu;WriterTxthat setsresolved_pathonly where it is still NULL and moves the watermark. The watermark lives in a one-row table,resolved_path_backfill_state.Restarts: a restart repeats at most the one uncommitted batch. Rows that still do not resolve stay NULL and are not retried in a later pass.
Readiness guard: the pass and every batch run only when three things hold. The edge build counts as complete only once a call has caught up, meaning it read fewer observations than its 50,000-row cap. The warm-up loops on rows scanned, not on edges produced. A full batch that yields no edge cannot move the builder's watermark; the warm-up then stops, logs it, and leaves the graph unpublished, so the backfill waits. Otherwise they write nothing, not even the watermark, and the background pass waits for the next post-build graph and resumes from the persisted watermark. The three conditions:
neighbor_edgesbuild that succeeded and caught up:StartNeighborEdgesBuilderpublishes it after such a warm-up, or after such a tick;Without the guard, a pass on the pre-build snapshot or on an empty graph would resolve only unique prefixes and move the watermark past every other row for good.
Config (defaults):
resolvedPathBackfill.batchSize= 500;resolvedPathBackfill.pauseMs= 250;resolvedPathBackfill.disabled= false.They are documented in
config.example.jsonanddocs/user-guide/configuration.md. ForbatchSizeandpauseMs, 0 or omitted means the default, so the pause cannot be turned off; the smallest pause is 1 ms. These are candidates for the customizer later (AGENTS.md rule 8).Server visibility: on master, the server's relay-credit indexes read
resolved_pathonly at start-up (Load); with fix(store): index live observations from the persisted resolved_path, as Load does (#158) #182 they also read it when a row is first polled. Either way a backfilled value reaches those indexes only on the server's next restart. The server is not changed. A possible follow-up is a periodic server-side re-read of rows whoseresolved_pathchanged, for example keyed by the backfill watermark.Replay on staging data (public API, read-only)
Staging build int-0b11401d. 3,000 transmissions in 10 windows over 44 h, with 73,275 observations that carry a path. The replay runs this branch's Go code. Two graphs:
/api/analytics/neighbor-graph;neighbor_edgesitself is not public;The old resolver reproduces staging's persisted NULL/non-NULL status for 98.5 % of rows.
resolved_pathNULLtraffic_share_score)Performance
Resolution per observation (
InsertTransmissionhot path): benchmark on the staging sample, one op = one observation, n = 6.That is +1.6 µs per observation; at staging's ~4–5 observations/s it is about 10 µs of CPU per second. No new maps outside the call, no work under locks, and no
map[string]interface{}.Backfill write-lock hold per 500-row batch:
TestResolvedPathBackfill_WriteHoldUnderBudget_188, is 250 ms.Tests
TestObserverAnchor_LastHopResolvesViaObserver_188TestObserverAnchor_BackwardWalkResolvesEarlierHops_188TestObserverAnchor_TwoNeighborCandidatesStayNil_188TestObserverAnchor_AdvertForwardResultUnchanged_188TestObserverAnchor_ForwardBackwardMeetOrNil_188(agree / disagree / same node twice)TestObserverAnchor_NeverContradictsForward_Random_188(2,000 random graphs)TestObserverAnchor_HashSizes_188(1/2/3-byte)TestObserverAnchor_DirectRouteIsNotAnchored_188TestObserverAnchor_UnknownObserverHasNoEffect_188TestInsertTransmission_ObserverAnchorResolvesLastHop_188TestIsRelayRole_MatchesServerCanAppearInPath_188TestBuildPrefixIndex_RelayRolesOnly_188TestInsertTransmission_CompanionNeverStoredAsRelay_188TestResolvedPathBackfill_ResolvesRowsIngestedBeforePriming_188…_BatchBoundAndWatermark_188…_Idempotent_188…_ResumesAfterRestart_188…_StopKeepsWatermark_188…_WriteHoldUnderBudget_188…_RefusesBeforePriming_188TestResolvedPathBackfillSettings_Defaults_188role = 'repeater'. Expectations are unchanged.Mutants (each in a copy of the tree, all red)
TestObserverAnchor_LastHopResolvesViaObserver_188TestObserverAnchor_TwoNeighborCandidatesStayNil_188TestBuildPrefixIndex_RelayRolesOnly_188,TestInsertTransmission_CompanionNeverStoredAsRelay_188TestResolvedPathBackfill_ResumesAfterRestart_188TestResolvedPathBackfill_BatchBoundAndWatermark_188TestResolvedPathBackfill_BatchBoundAndWatermark_188TestResolvedPathBackfill_RefusesBeforePriming_188TestObserverAnchor_ObserverIsNeverItsOwnLastHop_188…_BackwardWalkBreaksOnNilHop_188…_BackwardWalkExcludesResolvedHops_188from_pubkeyfor non-ADVERT rows (m12)…_FromPubkeyOnlyForAdverts_188TestObserverAnchor_ForwardBackwardMeetOrNil_188/disagree…TestObserverAnchor_DirectRouteIsNotAnchored_188The "priming after drain" mutant belongs to point 3, which is not in this PR.
Review follow-up (findings 1, 2, 4, 5, 8, and the second review's P2-1 and P2-2)
f9bae200(red),cbaa9ca6(fix),f5add678…_UsesGraphFromFirstEdgeBuild_188,…_EmptyGraphKeepsWatermark_188,…_StartWaitsForFirstEdgeBuild_188,TestResolvedPathBackfillReady_188,TestNeighborEdgesBuilder_TickPublishesBuiltGraphAfterFailedWarmUp_188pauseMs: 0means the defaultd1db161cTestConfigExample_ResolvedPathBackfill_188,TestResolvedPathBackfillSettings_ZeroMeansDefault_188e610dbb8(red),503ea758(fix)TestObserverAnchor_ObserverIsNeverItsOwnLastHop_188,TestObserverAnchor_ObserverMayBeAnEarlierHop_1880c623dba…_BackwardWalkBreaksOnNilHop_188,…_BackwardWalkExcludesResolvedHops_188,TestResolvedPathBackfill_FromPubkeyOnlyForAdverts_1882e2886afe369ae69(red),da0494c4(fix)TestNeighborEdgesBuilder_WarmUpScansPastFullBatchWithFewEdges_190,TestResolvedPathBackfill_WaitsWhileEdgeBuildCannotCatchUp_190addccb88(red),194512d2(fix),0f4c167aTestNeighborEdgesBuilder_ObserverEdgeOnlyForFloodRoutes_190The resolver benchmark is now in the repo:
BenchmarkResolveObservationPath_188(700 relays, 60 observers, 1–5 hop 1-byte flood paths). Finding 4 costs nothing measurable: median 3,875 → 3,890 ns/op, 119 B/op and 5 allocs/op on both sides (n = 12 each, interleaved, same machine).Deploy notes and further effects (added at merge, from the review of
0f4c167a)Two further effects of section 2b:
countnever goes away for node pairs that keep getting real flood observations. Neighbour-prune only removes edges with no new observations for 5 days, so such counts stay somewhat inflated rather than being corrected.Deploy:
[resolved_path_backfill] done, and only then read the traffic-share sum.Not verified
neighbor_edges.main.goordering: the background pass now waits for the first post-build graph whatever the call order (…_StartWaitsForFirstEdgeBuild_188).🤖 Generated with Claude Code