Skip to content

ingestor resolves only ~half of relay paths (resolved_path NULL in ~53 % of observations); possible hasResolvedPath startup race #184

Description

@dborup

Summary

Two related questions came up while investigating why traffic share still rises after #162. The investigation led to #182.

  1. The ingestor resolves only about half of the relay paths. On staging, only about 47 % of observations that carry a path have a non-null resolved_path. Since fix(store): index live observations from the persisted resolved_path, as Load does (#158) #182, the server indexes relay credit (byPathHop, traffic share, handleNodePaths, paths-through) only from that persisted value, for both live ingest and load. So roughly half of all relayed traffic gives no relay credit. Is that the intended level, or is the ingestor's resolver missing hops it could resolve?
  2. A possible startup race on hasResolvedPath. The store's startup load may run before the flag that tells it to select resolved_path is set.

Relates to #158, #162, #182.

1. Ingestor resolution coverage

Measurements (staging, read-only, 2026-10-03):

  • Per observation: in a sample of 500 non-grouped packets, 283 carry a path and 132 of those have a non-null resolved_path (46.6 %). Inside the non-null rows, only 48 of 1,106 hops are null (4.3 %). So it is mostly all-or-nothing per observation.
  • Per hop: over 8 minutes (2,860 non-advert observations, 21,651 hops), the server's own resolver resolved 98.5 % of hops and the ingestor's persisted resolution 59.9 %. That gives 54.3 vs 31.4 distinct relay pubkeys per transmission.

Where:

Questions:

  • Why is a whole observation's resolved_path NULL in about 53 % of cases? Possible causes:

    • an empty neighbor graph at ingestor start;
    • a missing anchor (fromPubkey);
    • prefix collisions that are common in this mesh;
    • an early return in the code;
    • something else.

    Measure each cause separately.

  • Is the conservative "exactly one adjacent candidate" rule right for scoring? Or should the ingestor resolve more hops, for example with a geo or affinity tie-break like the server's, while staying deterministic per transmission rather than per observer?

  • A row is written once. ON CONFLICT … COALESCE(excluded.resolved_path, resolved_path) only fills it on a duplicate delivery. Would a periodic backfill for NULL rows, once the graph has filled, close the gap? Note: such a backfill must live in the ingestor (bug(db): vacuumOnStartup fails with SQLITE_BUSY when ingestor + server share DB (auto_vacuum migration #919 broken in single-container topology) Kpa-clawbot/CoreScope#1283 read/write separation).

Deliverable: a measurement of the NULL causes on real data, using staging's public API or a staging-sized fixture, and a proposal. Changes to the resolver are a separate decision, because they change traffic share and relay attribution network-wide.

2. Possible startup race on hasResolvedPath

The #182 review found the following by reading the code; it is not reproduced:

  • In cmd/server/main.go, waitForDBSchema runs at ~200, and store.RunStartupLoad(chunkSize) starts in a goroutine at ~253.
  • database.hasResolvedPathFlag.forceTrue() comes later, at ~292. The comment there calls it an optimisation, because the flag self-heals through a PRAGMA re-probe every 2 s.
  • If the initial OpenDB probe saw the flag as false, the hot-window Load could run before forceTrue() or the healer. It then takes the NULL branch: byNode only, no byPathHop credit for that window, until the next restart.

This belongs to the known family of startup races around schema detection, where optional columns can be hidden.

Fix idea: set the flag before RunStartupLoad starts, after the schema wait has confirmed the column. Alternatively, have the store snapshot the flag once per load, after AssertReady.

Test: a startup test where the flag starts false and the column exists. Load must index byPathHop from resolved_path.

Acceptance

  • Part 1: a short report with measurements and a recommendation (no resolver change without a decision).
  • Part 2: a test that reproduces the race (or shows it cannot happen) and a fix if needed. The fork guards stay unchanged.

Activity

  1. dborup commented on Oct 3, 2026

    @dborup
    OwnerAuthor

    Part 2 (startup race on hasResolvedPath) is fixed by #186 (merged as ae595682).

    • The race in the form described above cannot happen: AssertReady requires observations.resolved_path, and the flags are re-probed after the schema wait.
    • A narrower form could happen: if the PRAGMA probe failed right after AssertReady, the first, newest chunk was loaded without resolved_path.
    • waitForDBSchema now sets the flag as soon as AssertReady passes, and the late forceTrue() in main.go is removed. A test covers this.

    Part 1 (why the ingestor resolves only about half of the relay paths): the investigation found that every NULL row is a non-advert observation with 1-byte hop hashes, because the ingestor has no anchor for those packets. The fix (the observer as anchor for the last hop, a relay-only prefix index, and a backfill) is tracked in #188, so this issue stays open until #188 lands.

    Not fixed by #186: Load/LoadChunked read the schema flags twice per chunk (once for the SELECT, once for the Scan); a flip in between drops the chunk's rows. #182 fixes the same pattern on the live-ingest paths only.

  2. dborup commented on Oct 3, 2026

    @dborup
    OwnerAuthor

    Status: #182 is merged (as 9d29daec). The server now indexes live observations from the ingestor's persisted resolved_path, the same way Load does, so the live and startup halves of the store agree.

    This does not change the ingestor's coverage. The ~53 % NULL resolved_path (part 1 of this issue) is addressed by #188 / #190, which is in progress. This issue stays open until that lands.

  3. dborup commented on Oct 4, 2026

    @dborup
    OwnerAuthor

    Fixed by #190 (merged as c7798e0b) together with #186 and #182:

    The effect on real data will be checked in the next staging round.

    Leftover: the double schema-flag read in Load/LoadChunked, noted earlier in this issue, is not addressed.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions