Repository navigation
Record every observed route per transmission and show mixed ADVERTs in Relay Airtime Share - #93
Merged
Merged
Conversation
Issue #89: a content hash can be observed on more than one raw route type (e.g. a contact re-shared as a zero-hop advert). route_type keeps the route of the first inserted observation; route_mask now records all of them, independent of ingest order. - internal/packetpath: canonical RouteTypeFromHeader, RouteTypeFromRawHex and RouteMaskBit plus the flood/direct mask constants; the ingestor decoder uses RouteTypeFromHeader - internal/dbschema: transmissions.route_mask (nullable INTEGER, bit r = raw route r) added by Apply as a metadata-only ALTER and asserted by AssertReady; existing rows stay NULL ("not backfilled"). The pending index is defined here but built later by the async backfill, not at start-up - ingestor: the first insert stores the bit of its own route (0 for an invalid route); a later observation ORs its bit into a known mask in SQL, after the observation insert. route_type, dedup and first_seen/last_seen behaviour are unchanged Tests cover every insert order of each scenario, same and different observers, redelivery, ingestor restart, 200 barrier-started parallel races, PATH on routes 0 and 1, invalid routes, and that legacy NULL rows are left alone and bits are never cleared. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rows written before transmissions.route_mask existed are NULL. An async migration (tx_route_mask_backfill_v1) fills them with the bit of their stored route_type plus every bit readable from surviving observations.raw_hex headers, parsed with packetpath.RouteTypeFromRawHex. The result is a lower bound: variants whose frames were overwritten are not invented. Unparseable or truncated frames are skipped. - the pending partial index idx_tx_route_mask_null is built inside the backfill. On a 4.95 GB staging copy that took 18-19.5 s cold, so it runs after MQTT subscribe and ingestBuffer.Ready(), where IngestBuffer absorbs the one-time write stall - batches of 100 rows, each in its own transaction under writerMu, with a 50 ms yield. Measured on the staging copy: hold p99 161 ms, a concurrent live write waited p99 121 ms, 9.3 min in total - only rows still NULL are updated, so a crash or cancel resumes where it stopped; shutdown cancels the run before disconnecting MQTT and it resumes on the next start - a completed migration is re-run when NULL rows reappear (rows inserted by an older ingestor after a rollback) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The server reads transmissions.route_mask in every load path (Load,
LoadChunked, loadChunk, IngestNewFromDB and IngestNewObservations) into
two StoreTx fields that fit the existing padding (StoreTx stays 320
bytes on 64-bit). IngestNewObservations joins the transmission's mask
and also ORs in the new observation's own route bit, so a poll that
lands between the ingestor's observation insert and its mask OR still
matches a later cold load. Un-backfilled (NULL) rows stay unknown.
advertRouteClass is the single classifier: only routes 0/1 -> flood,
only 2/3 -> zero_hop, both -> mixed. Without usable mask bits it falls
back to the legacy first-inserted route_type, else legacy. Relay Airtime
Share gains route_class "mixed" / "ADVERT (mixed)"; a mixed hash counts
once and its relays stay on that row. Payload Type Mix is unchanged.
route_mask_backfill {status: pending|backfilling|complete, remaining} is
reported on /api/healthz and in the Relay Airtime Share response. The
server computes it read-only from the database (complete only when no
row has route_mask IS NULL), cached for 30 s. OpenAPI describes mixed,
the fallback and the Time-on-Air caveat.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A row with route_class "mixed" (#89: the same payload was observed on both flood and zero-hop routes) gets a fixed theme colour, var(--status-purple), so it reads the same in any position, and a tooltip line saying so and that it is counted once. Identity stays data-payload-type/data-route-class from the API; other rows and older responses without mixed render as before. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Two gaps found while mutation testing #89: - the mixed row's colour and tooltip must come from route_class, not from a "(mixed)" display label (a relabelled mixed row keeps them, a flood row labelled "(mixed)" does not get them) - Payload Type Mix must stay one combined ADVERT entry when transmissions carry flood, zero-hop and mixed route masks; the existing test only used unknown masks, so a split keyed on the mask survived Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…isible The ingestor ORed an observation's route bit into transmissions.route_mask after writing the observation row, in separate autocommits. A server poll in between stored the observation with a mask that lacked its bit, and the server never re-read it: the live view could permanently disagree with a cold load. Run the OR before the observation insert. writerMu already keeps the backfill out of InsertTransmission, so the old ordering argument (and its comment) did not hold. The server-side mergeObservationRoute hedge covered the same gap for one path only and is no longer needed. A trigger-based test records the mask at the moment each observation row is written, including the upsert; it failed before this change. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The server read transmissions.route_mask only when it loaded a row or saw a new observation of it. The ingestor backfills legacy rows after both processes have started (about nine minutes on the staging copy), so a running server kept classifying those ADVERTs by their first-inserted route until a restart, while route_mask_backfill already said "complete". Transmissions loaded while their mask is still NULL are queued. On every poller tick the server re-reads up to 2,000 of them (read-only, IN-lists of 500 ids by primary key), in ascending id order like the backfill, and stops at the first id that is still NULL. Rows that are gone or evicted are dropped. Merged masks invalidate the RF cache that holds Relay Airtime Share. route_mask_backfill now reports "complete" only when the database has no NULL row and this server has read every backfilled mask; until then it reports "backfilling" with the number still to read. OpenAPI says so. Also fixes a stale file reference in db.go. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…iter The pending-index build is one long statement (about 18 s cold on the staging copy). It ran on the single writer connection without writerMu, so the stall showed up as wait time of the MQTT handler queued behind it. It now holds writerMu and is recorded as route_mask_index. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Checks that every rendered row carries the (type, route_class) identity and label of its API row and that route_mask_backfill reports a status, then injects a flood / mixed / zero-hop response via page.route and asserts the mixed row's identity, var(--status-purple) colour and tooltip. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…sible A second independent review found that the background chunk loader queued a NULL-mask transmission while scanning the chunk into local maps, before the transmission reached s.byTxID. A poller tick in that window found no transmission for the id, dropped it as evicted, and the transmission then stayed unknown for good while route_mask_backfill reported "complete". Server and ingestor start together, and the backfill (oldest ids up) meets the background loader (newest chunks down), so this is a normal deploy. loadChunk now queues an id in its s.mu merge section, when it publishes the transmission in s.byTxID. A test hook between scan and merge reproduces the window; the chunked start-up paths (LoadChunked, loadChunk) now have their own queueing tests. Also from the review: - The refresh is serialised by its own mutex, and load paths append to an unsorted inbox that the refresh sorts outside s.mu and the inbox lock. - Nothing is queued without the route_mask column (test databases only). - Relay Airtime Share reads the backfill status before computing its rows, so a refresh landing in between errs toward "backfilling". - OpenAPI: "remaining" is an upper bound once the database is complete. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
When an observation brought a route bit its transmission did not have yet, the route_mask OR ran as its own autocommit and a failure was only logged; the observation was still written. The mask was already known, so the backfill (route_mask IS NULL only) would never repair it: the transmission stayed permanently short of that route, and the server could show an observation without its bit. The OR and the observation insert/upsert now run in one transaction, only on that path. If either fails, neither becomes visible; the failure is counted once in WriteErrors and logged as a failed observation write, and InsertTransmission stays non-fatal as before. All statements in the transaction run on it: the store has a single connection. Every other path (new transmission, legacy NULL mask, bit already set) is unchanged. Test-only triggers make either the route_mask update or the observation write fail, for both a new observation row and an upsert, and check that mask and observations stay unchanged, the counters move once, and a later redelivery stores both. The test fails on the previous code. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The comments still said the ADVERT route class is the route of the first inserted observation. Describe the current aggregation: route bits 0/1 is flood, 2/3 is zero-hop (direct), both groups is mixed, each hash counted once; the legacy route_type fallback applies only while the mask is not known or has no usable bits. Comment-only. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A new route bit that reaches an existing transmission through an upsert of an existing observation row (same observer and path) keeps the observation id, so a running server, which polls new ids, never learns about it and keeps classifying the transmission as flood or zero-hop instead of mixed until a restart. route_mask_changes is the durable change log from the ingestor to running servers: one row per actual mask change, with the transmission's resulting full mask and the ingest time. AUTOINCREMENT keeps ids from being reused after deletes, because servers poll by id. It is a new, empty table plus an index on transmission_id: no scan or rewrite of existing data (6.6-6.8 ms + 3.2-4.0 ms cold on a 4.95 GB staging copy, 0 bytes of growth). AssertReady requires it, so the ingestor is deployed first, as for every schema step. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
When an observation brings a new route bit to an existing transmission, the OR update, the observation insert or upsert and, when the mask actually changed, one route_mask_changes row with the resulting mask now run in one transaction. Any failure rolls all of it back and is counted once; nothing is logged for a new transmission, a repeated route or a legacy NULL mask (the backfill owns those). Retention: PruneOldPackets deletes a batch's change rows in the same transaction as its transmissions, through idx_route_mask_changes_tx (plan asserted). PruneOrphanRouteMaskChanges removes rows whose transmission is gone by any other path, at start-up and after the daily packet prune. Each writer transaction examines a bounded id range: at the theoretical maximum of 1.49 M rows the writer hold is p99 36.6 ms (the first, LIMIT-on-delete version held it for 22 s). Rows of existing transmissions are never deleted and there is no time-based expiry. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…estart A same-observer, same-path observation with a new route updates the observation row in place: the database mask gains the bit, but the poller only reads new transmission and observation ids, so the in-memory mask and Relay Airtime Share stayed flood or zero-hop instead of mixed until a restart (reproduced deterministically before this change). Every poller tick (now Poller.pollStore, so tests drive the production path) reads route_mask_changes past a cursor in id order, at most 1,000 rows, without s.mu; ORs them per transmission and merges them under s.mu into transmissions the store holds, never re-creating one. The Relay Airtime Share cache is invalidated only when a mask actually changed. Changes for transmissions not in the store yet are parked and applied when the transmission is published; they are dropped once the start-up load is over (Load or RunStartupLoad returned, also after a partial background fill) and the store has seen past that id. The watermark MAX(id) is read before Load or LoadChunked reads any transmission, so changes committed during the load are applied by the poller; old rows are never re-read after a restart. If the watermark cannot be read, the log is read from the start rather than skipped. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The server and the ingestor start together under supervisord. The server required route_mask_changes and exited when it was missing; when the ingestor's start-up Apply was held up (a 70 s write lock in the staging gate), supervisord's back-off gave up after about 64.6 s and left the server FATAL for good. The server now waits in-process for the schema, bounded and cancellable: 250 ms back-off growing by 1.5x to 2 s, deadline 120 s, SIGINT/SIGTERM ends the wait with exit 0. It retries only a typed, transient *dbschema.NotReadyError: route_mask_changes not created yet (optionally with transmissions.route_mask, added by the same Apply), or SQLite BUSY/LOCKED by result code. Every other defect fails at once: other missing items, a malformed change log (columns, AUTOINCREMENT, exact index), a recorded marker without its table, permissions, corruption, closed handles and unknown errors. After the deadline the server exits non-zero after having run past startsecs, so supervisord restarts it without consuming start retries. The server still never writes the schema. AssertReady reads every probe in one read transaction, so a migration committing mid-check is never half-seen as permanent. The ingestor now creates the change log, its index and its marker in one transaction. After the wait, the server re-reads its optional-column flags so a pre-#93 database loads route_mask and queues rows for the backfill refresh. Readiness stays red while waiting: the HTTP listener starts only after the schema is verified. Tests use a fake clock and real SQLite (no sleeps), plus process tests that run the real main() in a child process. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Sep 25, 2026
Closed
6 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Background
Issue #89 is background only; this PR does not close it.
One content hash can be observed on more than one raw route. MeshCore firmware re-sends a stored advert as
ROUTE_TYPE_TRANSPORT_DIRECT(shareContactZeroHop), so the same ADVERT can arrive both flooded and zero-hop.transmissions.route_typestores only the first route that got inserted. Relay Airtime Share (PR #86) therefore put a transmission on the flood row or the zero-hop row depending on which observation happened to arrive first.Contract
A transmission stays one payload/content hash. The content hash still ignores route bits, transport codes and path.
A new column,
transmissions.route_mask, records every raw route value (0–3) observed for that transmission. Bitrstands for router.The mask is monotonic: every observation ORs its bit in, atomically in SQL, and no bit is ever cleared.
transmissions.route_typeis unchanged. It keeps its legacy meaning (the first-inserted route), and every other consumer keeps reading it as before.Relay Airtime Share classifies each ADVERT by its mask, and counts it once:
route_classmixedADVERT (mixed)floodADVERT (flood)zero_hopADVERT (zero-hop)A row whose mask is not known yet (legacy, or not backfilled yet) falls back to the previous
route_typeclassification.Payload Type Mix still shows ADVERT as one combined type.
The historical backfill is a lower bound. It never invents information that was lost.
Schema and migration
internal/dbschema.Apply(ingestor) only runsALTER TABLE transmissions ADD COLUMN route_mask INTEGERand writes the markertransmissions_route_mask_v1. There is no default, so existing rows stayNULL, which means "not backfilled".AssertReady(server) requires the column.The partial index
idx_tx_route_mask_null ON transmissions(id) WHERE route_mask IS NULLand the backfill run together as the async migrationtx_route_mask_backfill_v1.ingestBuffer.Ready(), so the IngestBuffer covers the one index build that can't be split into chunks.writerMuand is recorded as the writer componentroute_mask_index.The backfill runs in batches of 100 transmissions. Each batch is one write transaction under
writerMu, followed by a 50 ms yield.NULLrows appear again after completion (for example, an older ingestor wrote rows after a rollback), the next start runs it again.The server stays read-only (
mode=ro). It reportsroute_mask_backfill: {status, remaining}on/api/healthzand in the Relay Airtime Share response:pendingbackfillingcompleteroute_mask IS NULLand this server has read every backfilled mask. It is never derived from a marker.Change log (added after review, see "Live refresh" below):
internal/dbschema.Applyalso createsroute_mask_changes, gated by the markerroute_mask_changes_v1.AssertReadyrequires the table, so the ingestor must be deployed before the server, as for every other schema step.It is a new, empty table: no scan or rewrite of existing data.
AUTOINCREMENTkeeps ids from being reused after deletes, because servers poll by id.Lower bound
A legacy row gets two things:
route_type;observations.raw_hexheader.A route variant whose frame no longer exists anywhere is not reconstructed. An example is a zero-hop frame overwritten by the same-observer/same-path upsert. Unparseable or truncated frames, and routes outside 0..3, contribute nothing.
Cold-cache migration measurement
The budget was written down before measuring. The measurement ran on a consistent SQLite backup of the staging database, made with the SQLite backup API and verified by SHA-256.
--cpus 2 --memory 1g --network none.posix_fadvise(DONTNEED)on that run's own DB file, verified withfincore= 0.Apply(ALTER + marker), includingOpenStoreOpenStore44.5 msquick_checkokobservationsThe first design built the index inside
Apply, before subscribe. It failed the original start-up budget of ≤ 15 s, which is why the index and the backfill moved to the async migration after subscribe.On the copy, the final implementation left 0 rows
NULL. Its mask checksum matches an independent Python calculation. It found 21 mixed transmissions (8 TXT_MSG, 7 ACK, 6 ADVERT).Order, redelivery and parallel ingest
Every permutation of each scenario gives the same mask, and
route_typestays the first-inserted route. The scenarios:The mask survives an ingestor restart.
200 fresh payloads, each with both variants racing for the writer, all end with
1010.A route outside 0..3 sets no bit, but still writes a known mask of
0.PATH seen with and without transport codes keeps both flood bits.
Live ingest never writes into a legacy
NULLrow; that row belongs to the backfill.A backfill batch running at the same time as live ingest loses no bits, because both hold
writerMu.An observation row never becomes visible without its route bit. When an observation brings a new bit, three writes run in one SQLite transaction:
route_mask;route_mask_changesrow with the resulting mask.If any of them, or the commit, fails, the transaction rolls back and nothing becomes visible. The failure is counted once in
WriteErrors, andInsertTransmissionstays non-fatal as before. A trigger-based test records the mask at the moment each observation row is written, including the upsert. Fault-injection tests make each of the three writes fail, for both a new row and an upsert.No change row is written for a new transmission (the server reads its full mask when it loads the row), a repeated route, or a legacy
NULLmask (the backfill owns it).Server and in-memory model
StoreTxgetsrouteMask uint8androuteMaskKnown boolin existing padding.unsafe.Sizeof(StoreTx{})is still 320, and a test pins it.Every load path reads
t.route_mask:Load,LoadChunked/loadChunk,IngestNewFromDBandIngestNewObservations. Tests assert that the incremental view equals a cold load.Backfilled masks reach a running server. Transmissions loaded while their mask is still
NULLare queued once they are visible in the store. That includes the background chunk loader, which queues in its locked merge step. On every poller tick the server re-reads up to 2,000 of them:NULL.Rows that are gone or evicted are dropped. Merged masks invalidate the RF cache that holds Relay Airtime Share.
The refresh is serialised by its own mutex. Load paths append to an unsorted inbox, which the refresh sorts outside
s.mu.advertRouteClassinrelay_airtime_share.gois the only classifier.Live refresh of same-row upserts
The bug, reproduced deterministically before the fix: the same observer hears the same path on a new route. The observation is upserted in place, so it keeps its id and no new row is created, and the database gets both bits (
1001). But the poller only reads new transmission and observation ids. After three production poll ticks the in-memory mask was still0001or1000, and Relay Airtime Share still showedADVERT (flood)orADVERT (zero-hop)instead ofADVERT (mixed), until a restart.The fix: the server reads
route_mask_changeson every poller tick. The tick body is extracted intoPoller.pollStore, so tests drive the exact production path.s.mu. Masks for the same transmission are ORed together.s.mu, and only to transmissions the store already holds. A transmission is never re-created. The Relay Airtime Share cache is invalidated only when a mask actually changed.LoadorRunStartupLoadreturned, whether it succeeded, failed or loaded only part) and the store has seen past that id.Start-up: the watermark
MAX(id)is read beforeLoadorLoadChunkedreads any transmission.transmissions.route_maskis authoritative.Tests cover:
LoadandLoadChunked), and before the first tick;Retention of the change log
PruneOldPacketsdeletes a batch's change rows in the same transaction as its transmissions, throughidx_route_mask_changes_tx(the query plan is asserted in a test).PruneOrphanRouteMaskChangesremoves rows whose transmission is gone by any other path. It runs at start-up and after the daily packet prune. Each writer transaction examines a bounded id range, so its work does not depend on how far apart orphans are. It never deletes a row of an existing transmission, and there is no time-based expiry, so a slow server cannot miss a change of a transmission it still holds.API and frontend
The JSON changes are additive only:
route_classvalue,"mixed";route_mask_backfillobject.OpenAPI documents both, plus the route-mask semantics and the fallback.
Relay Airtime Share renders the mixed row with:
data-route-class="mixed";var(--status-purple), whatever its position;The colour and the tooltip come from
route_class, not from the label.Browser verified: local scratch server, desktop 1280 px and 375 px (details below; not staging).
E2E assertion added: test-e2e-playwright.js:855
Browser validation (local scratch datasets)
The branch server ran against these datasets:
pending(index missing) andbackfilling.Every set also had non-ADVERT traffic and two
UNKrows. For each set, a script compared every API row with its DOM row:data-payload-typeanddata-route-class;cnt %/air %text;rgb(168, 85, 247)Master's frontend, served against the branch API as a negative control, renders the mixed row as a plain row without the purple or the tooltip. That shows the check detects the change, and that the older frontend degrades cleanly.
At 375 px the page has no horizontal scroll. Inside the chart, the fixed 80/180 px grid overflows by 38 px. That is identical with master's frontend: a baseline, not changed here.
The new Playwright test checks API↔DOM identity on the fixture. It then injects a flood / mixed / zero-hop response through
page.routeand asserts the mixed row's identity, colour and tooltip. Its in-page assertions were dry-run against the scratch server. The full Playwright suite was not run locally, because no Playwright module is installed here; CI runs it.Start-up schema wait (staging gate blocker)
Problem reproduced
The server and the ingestor start together under the unchanged supervisord config (
autorestart=true,startretries=10,startsecs=2). On a database without this PR's change log, the server used to calllog.Fatalfon the first missingroute_mask_changes. In the isolated staging gate, a 70 s write lock held the ingestor's start-upApplyback. supervisord restarted the server with growing back-off and gave up after about 64.6 s, leaving itFATAL. It stayed stopped even though the ingestor created the table a few seconds later.Start-up contract
dbschema.AssertReadyreturns a typed*dbschema.NotReadyError. Only these are retried (Transient()):route_mask_changesdoes not exist yet and its marker is not recorded;transmissions.route_maskis missing, because the sameApplyadds it just before the table;BUSY/LOCKED. This is classified by the result code (Code() & 0xff∈ {5, 6}), never by error text.log.Fatalfas before:AUTOINCREMENT, or the index not exactly(transmission_id));route_mask_changes_v1recorded while the table is gone, since the ingestor never recreates it;AssertReadyruns every probe inside one read transaction. A migration committing mid-check is therefore seen either as not started or as done, never half-way (which would be misread as a permanent defect)./api/healthz) and the packet store load only start after the schema is verified. The optional pprof listener (ENABLE_PPROF, off by default) is unchanged. A process test asserts that the port refuses connections at the first wait log and at a status line.route_maskand queues NULL rows for the backfill refresh.shutdown requested while waiting for the ingestor's schema; exitingand exits 0; this is never reported as a timeout.schema still not ready after … (N attempts): …. By then it has run far longer thanstartsecs=2, so supervisord counts it as a started process and restarts it without using upstartretries. A slow migration therefore gives repeated 120 s waits, neverFATAL.mode=ro, and the ingestor remains the only schema writer. supervisord, the Dockerfile, the entrypoint, workflows and deploy scripts are unchanged.Measured on the real image and supervisord (local Docker, isolated volume)
Image
corescope:pr93-local-final, built from4db9716fwith the repository's own Dockerfile. It runs the unchanged supervisord config and entrypoint, withDISABLE_CADDY=trueand--network none, on a copy of the E2E fixture reset to the pre-#93 state (noroute_maskcolumn, no change log, no markers). "Hold" means a separate process holdsBEGIN IMMEDIATEon the database from before the container starts, so the ingestor'sApplycannot run. Readiness is/api/healthzready:true, polled once per second.FATALschema ready after 253–255ms (2 attempts)in all 20schema ready after 9.239s (9 attempts)FATAL)schema ready after 1m21.552s (45 attempts)schema ready after 1m55.692s (62 attempts)schema ready after 3.329s (6 attempts)schema ready after 7.255s (8 attempts)docker stopduring a 60 s holdshutdown requested while waiting…;stopped: corescope-server (exit status 0)in 1.04 sdatabase is lockedafter its own 5 s busy timeout, and supervisord restarts it. It ran paststartsecseach time, so it never reachedFATALeither, even in the 250 s run.route_mask_changeswas created once;write path ready; draining backlog (0 dropped);listening onappeared exactly once per successful server process.gave up: corescope-server entered FATAL stateat 64.6 s in the isolated staging gate.Limitations of the wait
_busy_timeout=5000(_journal_mode=WALlikewise; WAL persists in the file from the ingestor). A probe that meets a writer's lock therefore reportsBUSYat once, and the wait retries it. This PR does not change the DSN.signal.Notify(store load, a few seconds) takes Go's default action. That window is unchanged by this PR, apart from a sub-millisecond window around the flag re-read. SIGTERM during the wait itself is handled.main. It is covered by a unit test onwaitForDBSchema, which runs the real start-up load and queue.Applyfor good, gives a server that restarts every 120 s with a clear log. It does not giveFATAL. The ingestor's own logs andstartretriesstill surface the ingestor failure.Independent review
Fresh reviewers, none of whom wrote the code, reviewed each step. Every blocker and should-fix they reported is fixed test-first in its own commit.
Review 1:
Review 2, blocker: the background chunk loader queued an id before its transmission was visible. Fixed: queued in the merge step.
Review 3: the route bit and the observation write are now atomic. Approved; its test gap (the observation write failing after the OR) was closed.
Review 4, design B, no blocker. Its should-fix items were fixed:
Its nits were also closed: the no-change guard and the orphan-prune batch bound are now tested, and the table DDL is shared with the tests.
Review 5 (the review-4 fixes): no blocker and no should-fix. Each of those fixes was killed by its own test under mutation, and the new tests passed 50 runs under
-race. Both of its nits were applied: the test hook's comment now says it runs unders.mu, and the orphan prune guards against a batch size ≤ 0.Remaining pre-existing gap: none for same-row upserts; that live-view limitation is fixed by the change log.
Review 6 (the start-up wait):
route_maskswitched off for its start-up load;Also fixed: probe errors are now structured, not text; the index must be exactly
(transmission_id); and the log wording.Review 7 (the review-6 fixes): all the fixes were confirmed.
AssertReadyread the schema across several snapshots, so anApplycommitting mid-check could be misread as permanent. The reviewer reproduced this 71 and 104 times in 1,500 runs. It is fixed with one read transaction and a deterministic test that commits between two probes.Mutation: all 18 mutants of the wait and the classification are killed.
Mutations, race and performance
Mutations: 49 of 49 mutants killed, every kill naming a failing test (the runners stop on a build failure instead of counting it). The 18 change-log mutants cover:
LoadorLoadChunked, lazily, or reset to the end after a read error;RunStartupLoad;The earlier 31 mutants cover:
complete, from a marker or from masks the server has not read;NULLmerged as known;NULLas resolved, or missing from the poller;LoadChunkednot queueing, and queueing without the column;Race:
go test -race -count=10passes on the route-mask, change-log, backfill, prune, relay, healthz, cache, chunked-load and poller tests incmd/ingestorandcmd/server. The new focused tests pass 20 times in a row, and no test processes are left behind.Change log, measured:
CREATE TABLE+ index, cold, 4.95 GB staging copy (demo host,--cpus 2)OpenStoreincludingApply39–48 msb674e1e2)INSERT)Ingest hot path: the only extra work is one guarded
UPDATE … SET route_mask = route_mask | ?.Ingest benchmark: 15,000 tx × 21 observers, 1 in 500 mixed, measured before the OR moved ahead of the observation insert. The statements are the same.
--cpus 2, 2 alternating roundsServer poller refresh: at 500,000 queued transmissions, warm, locally:
s.mu)Server load:
StoreTxstays 320 bytes on both. Load time (3.6–4.0 s) and heap after load (206–213 MB) are equal within noise.Relay Airtime Share: cold compute is 7–15 ms (base) vs 17–21 ms (branch) on 15,000 packets; warm results are cached.
Known limitations and follow-ups (not in this PR)
s.mu. With the realistic 22-row log this costs nothing; at the theoretical 1.49 M rows it would be expensive for a while. It only happens after a failed watermark read.route_mask_changes(AssertReady). When both start together, it now waits for the ingestor's migration (see "Start-up schema wait"); the ingestor must still be deployed with or before the server. An older ingestor after a rollback writes no change rows, so a running server misses same-row bits again until its next restart; the database mask stays correct.Commit, not blocking.Rollbackis deferred right after a successfulBegin, and a failedCommitis returned, counted once inWriteErrorsand logged. On a failedCommit,database/sqlreturns the connection to the pool, andmodernc.org/sqlitev1.34.5 only checks it for an interrupt. A transaction could therefore stay open only if COMMIT failed without SQLite rolling back. In WAL mode, with the write lock already held and the only foreign key checked immediately, that needs an I/O or disk-full error. It could not be reproduced, and every other ingestor transaction has the same exposure. A possible follow-up is an explicitROLLBACKafter a failed commit.route_typeconsumers. Other consumers ofroute_type(for example packet list filters and other analytics) keep first-inserted semantics. Whether any of them should move toroute_maskneeds a separate decision.NULLrows.backfillinguntil the ingestor restarts and resumes. After a transient failure, it readspendinguntil the next start. It never readscompleteearly.--status-purpleis close to the palette violet used positionally by other rows. The labels distinguish them.cmd/server/routes.gostill has anINSERT INTO transmissionsthat fails on the read-only connection. This PR does not add it or change it.Preflight overrides
ALTER TABLE transmissions ADD COLUMN route_mask INTEGERis a nullable column without a default, so it is metadata-only in SQLite: no row rewrite, 8–15 ms on the 4.95 GB staging copy. It is annotatedPREFLIGHT: async=falsewith that reason. The partial index runs inRunAsyncMigrationwith anasync=trueannotation.Scope and deployment
4db9716fdoes not contain the wait and must be rebuilt from this head.🤖 Generated with Claude Code