Skip to content

Latest commit

 

History

History
948 lines (884 loc) · 59.5 KB

File metadata and controls

948 lines (884 loc) · 59.5 KB

Status

The one place "where is this project right now" lives. Replaces status prose that used to be duplicated (and drifting) across CLAUDE.md, PROGRESS.md, and README.md. For how we got here, see docs/history/CHANGELOG.md; for active V2 workstream gates, see ROADMAP.md; for the completed v1 phase record, see ROADMAP_V1.md.

Current version

v1.1.0 (src/version.h) — Workstream 12 closed 2026-09-07, the first stable minor of the V2 line. Its one outstanding item is stated rather than folded away: Portland field validation is deferred by explicit operator decision, not complete. Everything else in W12's gate is closed, including the WiFi off/on resource matrix (covered by the 8-hour soak) and the coverage thresholds (selected from an 80-pass campaign). Same convention as Phase 7's soak criterion, which was waived for a cycle and documented as waived rather than quietly satisfied.

Previously v1.0.7 (src/version.h). In the completed v1 series, MAJOR.MINOR tracks the build-order phase reached, not the phase in progress — see ROADMAP.md's Versioning section for the V2 workstream policy. Phase 9 (ENERGY_SWEEP/"Sweep") reached 2026-09-03: all five ROADMAP.md exit criteria closed, including two full 8-hour endurance soaks that caught and fixed a real logger_task stack overflow and a proactive radio_task stack margin fix — see "What's hardware-verified" below for the full writeup. v0.8.6-v0.8.9 were all PATCH bumps for out-of-sequence Phase 9/11 additions landed while Phase 9 itself was still open (Cell's repeat mode, FCC A/B block markers, the Sweep region setting) — see ROADMAP.md's Versioning section for why those stayed PATCH rather than MINOR. Phase 10 (Field Analyzer) reached MINOR status 2026-09-03 and closed all five exit criteria the same day — see the Phase 10 entry below. v0.10.1 is a same-day PATCH fixing a real bug the worst-case run itself surfaced (a repeat-mode Waterfall race condition); v0.10.2/v0.10.3 are same-day PATCHes for real, hardware-confirmed UI polish (Waterfall's frequency axis + repeat toggle, Meter's bar gauge/SNR/ channel-param block) — see the Phase 10 entry's "Post-closure UI polish" paragraph below. Same PATCH-not-MINOR convention as the Phase 9/11 bumps above throughout. Promoted to v1.0.0 the same day: ROADMAP.md's own documented gate for this promotion was Phase 10 closing, and only that — that's done. Phase 11 (Cell) was never part of the gate (added out of sequence, outside the original four-profile scope); its own two open items (below) are known, tracked gaps post-v1.0, an explicit operator call, not an oversight.

v1.0.7 is the stable V2 planning baseline. It is an operator-navigation PATCH release (Tools and Analyze moved into the menu and Activity became the read-only bounded-action status page); it does not change any completed phase gate or the two open Cell evidence items.

V2 entry state

The V2 product direction is adopted for planning, with its canonical workstream and gate policy in docs/ROADMAP.md and the detailed design in docs/research/V2_DESIGN.md. Workstream 12 (Survey truth) closed in v1.1.0. It ships a bounded one-bin Focus action and coverage labels. Portland field validation remains deferred, and the radio-away policy exception is limited to explicit single passes. The design-entry and acceptance record preserves the staged measurements below; descriptions of the bench-only first slice are historical, not current feature status.

The bench image now has one bounded Core-1 Focus request and a Core-0 focus.csv writer. Two paired 500 ms/eight-sample smoke checks at US Sweep bin 43 (912.750 MHz) observed P90/peak at -101.0/-101.0 dBm while quiet and -82.0/-82.0 then -87.0/-87.0 dBm during capped controlled 912.8125 MHz Heltec pulses—a 14--19 dB rise. Every request restored home listening and durably wrote its row. This is only fixture/transport evidence, not a calibration, coverage label, or activity claim. The repeatable bench harness is scripts/phase12_focus_bench.py; its pulse path requires explicit --with-pulse --allow-transmit.

The non-transmitting Focus behavior fixture also proved cancelled and injected failed rows restore before their durable write, and proved two-way refusal with Probe and Sweep.

A Focus request is now bounded in wall-clock time as well as in samples: past a dwell-plus-slack deadline it stops sampling and terminates as timeout, restoring home like every other exit path. That path and two-way Cell/Scope refusal are hardware-verified as of 2026-09-04 on bench build 3e31daa-dirty. An injected 1,500 ms sample-loop stall against a 100 ms request (1,100 ms deadline) produced a timeout row with successful home restore and 1,573 ms radio-away time; Focus and Cell, then Focus and Scope, each refused the other while it owned RX. Six terminal requests wrote six durable rows with zero queue or row drops. Reaching those paths on demand needs bench-image-only entrypoints (the stall is BENCH_FOCUS's optional fourth field; BENCH_ACTION starts and cancels the otherwise menu-only Cell and Scope); production rejects both. No coverage/activity claim is enabled.

The §6.2 controlled dwell matrix ran on 2026-09-04: 900 trials (5 positions x 3 dwell arms x 30 source-on and 30 source-off) in 54.7 min with zero transport errors, zero queue or row drops, and successful home restore on every trial. 14 of 15 arms separated a controlled source from ambient completely (30/30 vs 0/30). Its most consequential result is a limit, not a capability: detection tracks the source's airtime against Focus's sample spacing, not dwell length — with a fixed 8 samples a 2000 ms dwell observes eight instants, not 2000 ms, and one arm with a 94 ms source stopped separating there entirely. Bin-center offset was not the dominant term. Radio-away measured dwell plus ~73 ms, worst case 2,139 ms. No single fixed RSSI condition separates every arm, so the qualifying condition has a candidate (p90 >= -90 dBm) but not an accepted constant, and the coverage thresholds remain unselected. Full location-redacted summary: hardware-results/2026-09-04-phase12-focus-matrix.md. The §6.3 Watch-opportunity comparison has tooling but has not run, so Focus's maximum radio-away budget is still unapproved.

The qualifying RSSI condition was then measured at a realistic signal level (transmitter outdoors on the WiFi control bridge, receiver antenna fitted: source peak ~-85 dBm against a -99/-100 dBm ambient floor) and rejected. The p90 >= -90 dBm candidate came from a source ~70 dB hot; across 120 trials it fails, as does every floor-relative variant — one of ten metric/position combinations separates, by 2 dB, which is inside ordinary RSSI variance. An RSSI summary statistic cannot carry an activity claim at field levels. Coverage reporting is unaffected: valid passes, observation time and the RSSI summary stay honest. coverage stays blank, the activity count stays unpopulated, and the remaining routes are a per-pass count above an adaptive floor (qualifying_count, already reserved in the schema) or CAD/packet evidence.

The §6.3 Watch-opportunity comparison also completed: two 240 s arms against one independently timed reference train. Watch alone received 0.883 of the train; with Focus interleaved at a 48.1% away fraction it received 0.463, with non-overlapping 95% intervals. The loss is proportional to away time and nothing more (0.459 predicted against 0.463 measured), so Focus's recorded radio-away duration is an honest proxy for what a request costs Watch — no hidden retune or recovery penalty — and equally, there is no mitigation: half the time away is half the packets. Time between restored Watch windows ran to a 7.45 s maximum. The maximum away-time budget is deliberately still unapproved; the measurement supplies the exchange rate, not the policy.

A per-pass count above an adaptive floor then succeeded where the summary statistics failed. Counting samples at or above the pass's own median plus 6 dB, a threshold of two detected 57/60 source-on trials (95% CI [0.863, 0.983]) at the same field-level configuration, with a single flagged control trial that read -63 dBm against a -101 dBm median — a real transmission that was not ours, so the 1.7% false rate is an upper bound. It is a candidate, not a constant.

Two further sweeps (480 trials) then bounded it. The rule is a fraction of accepted samples (~4-5%), transferring between 500 ms and 2,000 ms passes, but no threshold works at 100 ms — six samples cannot both catch the source and reject ambient, so a short pass may report coverage and must not report activity. Detection also appeared to collapse below ~40% occupancy — but a second baseline at ~25 dB SNR (600 trials) withdrew that conclusion: 13 dB of link improvement took detection at 28.6% occupancy from 37% to 87%, so the cliff belonged to the first link, not to the instrument. Detection depends on occupancy and SNR together, and single-packet detection is untested rather than excluded. The 100 ms floor did survive both links (13 dB bought almost nothing there), as did the fraction-of-samples form.

That open question is now answered: one armed 148 ms packet inside a 2,000 ms pass — 7.4% occupancy — was detected 29/30 with 0/30 false positives at the stronger link. An RSSI-sampling pass can support a packet-level activity signal, but not unconditionally: the same rule at the weaker link failed on sources occupying four times as much of the pass, and a position carrying real traffic measured worse rather than better (a -57 dBm event in a control trial both raised false positives and suppressed counts by lifting the median). So Focus is least reliable where a band is busiest, and detection depends on link quality the device cannot know. qualifying_count stays unpopulated: wording an activity indication that stays true under those conditions is now a product decision rather than an open measurement.

The existing Phase 11 Cell feature remains partially hardware-verified and visible in "What's still open." It is deliberately scheduled as V2 Workstream 16 — Cell closeout, an optional post-core-V2 bonus: it does not block Workstreams 12–15 or the V2.0 composition release.

What's hardware-verified

Card views (5f472c4, confirmed on hardware 2026-09-06). The main carousel's five cards each carry an ordered list of views, cycled with up/down while left/right still moves the carousel: Meter/Scope under Radio, Sweep/Waterfall/Focus under Activity, Captures/Nodes/Probe under Channel, Cell under GPS. Enter runs the bounded action that refreshes whichever view is showing. Because every former Tools/Analyze page is now a card view, both of those menu groups were removed and the menu root is Profile / Trace / System; Trace returned to a root row. Flashed to the Cardputer-Adv and operator-confirmed: boots clean, IO expander P0 high, GPS fix sets the clock, radio listening on Core 1, ~257KB free heap after task start.

Two things were caught by the flash itself and fixed: the boot banner still described the pre-card-view key bindings (and named the wrong carousel keys — the pair is , and /, with ; and . as up/down), and Waterfall's empty state had been telling operators "Enter: start repeat Sweep" for two commits after Enter became single-shot. That second one is why key hints are now generated from cardSelectAction()/cardRepeatAction() rather than typed per page (cardHintLine()), and why the four post-hold IDLE headlines became the real terminal word dimmed plus a result age — IDLE described the radio while the view below it still showed a real result, so one card could disagree with itself about whether a sweep had happened.

Every card rebuilt in that language (691443c, flashed and walked through 2026-09-06). Radio became a six-value grid (RX/CRC/MISS over QUEUE/LOG/DROP — the radio layer over the storage pipeline); Channel a band map whose ticks mark where packets actually decoded and where Sweep found peaks, under a restored frequency hero; GPS a 60s satellites-used trace with a TAGGED card; System two trend lanes for heap and battery. Probe, Sweep and Cell moved onto the layout Focus had grown, each now leading with its finding rather than a state word, and Probe was promoted to Channel's view 2. Meter took the card shape.

The chrome converged along the way: a plot well (rails plus a floor, with the header hairline as its top edge and the floor as the plot baseline), cards widened to 78px from x=0 so they align with it, and WiFi moved to a fourth header status dot.

Three design rules came out of it and are worth keeping:

  • The mark follows the data type. Counts get bars, levels get lines or areas. Three of five bands had converged on the same 30 green bars and were indistinguishable; separating them was a chart-choice fix, not decoration.
  • Chrome follows what the band is. A well around something already bounded is doubling, which is why Radio's cells and Channel's axis lost their boxes.
  • Trends auto-scale with a minimum span. Pure min/max amplifies a 2KB heap wobble to full height; with a floor on the span, flat reads flat.

New state added: loggerRowsUntagged() (detection rows written without a fresh position — the wardriving quality number, counted at batch-accept and stated against detections seen), a 30-byte satellites ring and a 60-byte heap/battery ring. RAM 18.1%, flash 31.7%.

Still unverified for all of the above: no soak, no measurement of whether the extra per-card draw paths affect redraw timing under a running sweep, and loggerRowsUntagged() has never been exercised against a real GPS dropout — which is exactly the case GPS's TAGGED card exists to report.

Phases 0-9 are complete and hardware-verified: radio bring-up (Phase 1), the task/queue architecture + GPS + SD logging that makes up MVP-Beta (Phase 2), the WiFi AP + web command center (Phase 3), the MeshCore profile and live profile switch (Phase 4), the on-device menu UI (Phase 5), the UI architecture redesign (Phase 6), measured heap/stack budgets and soak (Phase 7 — its strict same-build repetition criterion was explicitly waived for that cycle, see docs/history/PROGRESS.md), bounded radio-owned discovery scanning with a source-backed candidate plan (Phase 8, DISCOVERY_SWEEP / "Probe"), and frequency-binned energy acquisition with selective Pass-B CAD (Phase 9, ENERGY_SWEEP / "Sweep", reached 2026-09-03 — full writeup below). Phase 8's statistical CAD false/miss matrix remains an explicit lab follow-up — the available bench can't provide a known-quiet RF control.

Phase 9 (ENERGY_SWEEP / "Sweep") is complete — all five exit criteria closed:

  • Pass A (frequency-binned energy acquisition across 868-923MHz) is hardware-verified end-to-end: real energy.csv peak rows spanning the full band, zero queue drops, clean home-restore across repeated sweeps.
  • Serial Control (SWEEP_START/SWEEP_CANCEL/STATUS) and a dedicated on-device Sweep result card are wired and hardware-verified.
  • The noise-floor margin is calibrated from a bench-run matrix: the shipped default moved from a 10.0dB placeholder to a measured 35.0dB, re-verified at 0/221 false peaks across three consecutive production sweeps.
  • Region setting (v0.8.9, added and hardware-verified 2026-09-01): System > Region narrows Sweep's scanned band to 902-923MHz (US, the default) instead of the full 868-923MHz range (Global), roughly halving scan time — 47 CFR § 15.247 cited in energy_plan.h. Persisted to /loratrace/region.txt. Cell and channel_plans.h are explicitly NOT region-aware (see docs/ROADMAP.md's Phase 9 follow-up bullets). Confirmed on real hardware (b7c845a): System > Region shows US by default and cycles to Global and back; a US-region Sweep shows the frequency bar reading 902/923 and completes noticeably faster; a Global-region Sweep reads 868/923 at full duration, matching pre-change behavior; the choice survives a reboot. One bug was caught and fixed during this verification: the System menu's row count was still hardcoded to 3 after Region became its 4th row, silently hiding the new row — ui_task.cpp's SYSTEM_GROUP_ITEMS GROUP entry now correctly says 4.
  • Pass B (CAD at peaks) is implemented and hardware-verified (landed 8367b73/5dae6ab, 2026-08-28/29 — this bullet was previously stale and said "has not started"). CAD runs immediately at each Pass-A peak, capped at PASS_B_MAX_PEAKS_PER_SWEEP (8) peaks per sweep, across the 10-row sourced SF/BW table in pass_b_plan.h. A CAD hit that promotes to a real packet logs as a Detection with off_grid = true (detectionClassification() returns unknown_lora_candidate, never a mission-profile name — the DESIGN.md §7.2 requirement that Pass B must not mislabel an off-grid hit as Reticulum). Per-combo confidence (PassBConfidence: STRONG/NOISY/UNVERIFIED) is a descriptive energy.csv column derived from a pooled 1,200-cycle bench matrix across three physical setups (open room, both radios shielded, Cardputer-only shielded) — only SF8/BW125 (STRONG, 2/60 quiet FP, 20/20 real-pulse detection) and SF11/BW500 (NOISY, 22/60 quiet FP, 19-20/20 detection) are reproducible enough to carry a confidence tag; the other eight combos stay UNVERIFIED pending more bench cycles. Full history, the SF-vs-time confound investigation, and the shielded-box findings are in docs/research/phase9-sweep-pass-b-design.md.
  • 923MHz-edge front-end rolloff — closed, no rolloff found (2026-09-01). ROADMAP.md's blocking unknown is resolved. Two lines of evidence, both real hardware:
    • Passive floor pass (BENCH_SWEEP_FLOOR, scripts/phase9_rolloff_bench.py): three real GLOBAL-band sweeps compared the mean floor in the top 5MHz near the front end's 923MHz ceiling against the rest of the band — +0.26dB, +0.52dB, -0.06dB, all inside the mid-band's own bin-to-bin spread. No signature, but passive (no transmission), so it couldn't rule out reduced gain on a real signal.
    • Injected-carrier pass (BENCH_RSSI_WINDOW, scripts/phase9_edge_carrier_bench.py): parks the radio at one fixed frequency and samples RSSI continuously for ~2s (a full Sweep's per-bin dwell, tens of ms, is too short to reliably coincide with an independently-timed transmitter's burst — a first attempt using a full Sweep came back a meaningless null result for exactly that reason). With the Heltec and Cardputer on matched 915MHz whip antennas and physically separated (desktop/under-desk, ruling out the near-field USB/clock coupling the Pass B shielded-box study already found raises apparent noise between close-together radios), 26 combined trials across two sessions at a mid-band (912.8125MHz, LONG_MODERATE) and an edge-band (920.625MHz, SHORT_SLOW) sourced candidate: mid-band captured the real signal 12/14 tries at -34.7dBm average, edge-band 14/14 at -36.9dBm average — only a 2.3dB gap, and the edge band was if anything more reliable, not less. An earlier close-together run showed a much bigger apparent gap (~25dB) that turned out to be a measurement artifact: computing signal "rise" as pulse-RSSI-minus-quiet-RSSI silently absorbs whatever the quiet baseline itself is doing, and the quiet baseline at 920.625MHz measured a real, reproducible ~24dB higher than at 912.8125MHz in this room with nothing transmitting — a genuine RF-environment fact (this is a dense urban area with AMI smart-meter deployments, which commonly use FHSS in the 900-928MHz band, a very plausible source), not a receiver characteristic. The fix was comparing absolute captured-signal strength instead of a delta from a baseline that isn't equal between the two frequencies being compared.
  • RTL-SDR ground truth for Pass B's CAD-at-arbitrary-bin question (2026-09-02). bench/rtl-sdr/sync_cad_capture.py triggers BENCH_PASS_B_CAD and a synchronized RTL-SDR capture at the same 918.5MHz test point, reading the raw result back via a new BENCH_PASS_B_CAD_RESULT opcode. All 10 PASS_B_SF_BW_CANDIDATES combos, quiet condition, n=5 each: every CAD_DETECTED (6 total instances, concentrated in SF11/BW500 as expected) showed a completely flat SDR waterfall at 918.5MHz — direct, not statistical, confirmation of the symbol-duration false-trigger hypothesis (docs/research/phase9-sweep-pass-b-design.md). A positive-control pass (--pulse, arming the Heltec once at 0ms delay immediately before each trigger — the timing this script's first, unsuccessful attempt got wrong) matched the original bench matrix exactly: the exact-match combo and the three whose bandwidth is a superset of the injected 125kHz signal all hit 5/5, confirmed independently on the SDR waterfall; everything else 0/5.
  • Three more Phase 9 exit criteria closed (2026-09-02):
    • Timing and home-away duration measured. Three real US-region sweeps: device-measured away time (EA, a new STATUS field exposing radioEnergyLastAwayMs(), previously internal-only) 3386-3438ms, home channel correctly restored every time.

    • Quiet-band behavior characterized with WiFi off/on. Three matched off/on pairs (WIFI_SET, a new Serial Control opcode mirroring the on-device menu toggle): zero peaks in every sweep regardless, and no meaningful timing difference (EA within ~50ms either way). WiFi does not introduce false Sweep peaks or measurably change sweep duration.

    • CAD never promotes energy alone to LoRa. 10 real BENCH_PASS_B_CAD attempts at the noisiest known combo (SF11/BW500): 5 came back CAD_DETECTED, and the real-packet-promotion counter (PBD) never moved once, across any attempt. Empirical confirmation, not just the code-level guarantee (off_grid is only ever set after a real decoded packet).

    • Injected low/mid/high carriers land in the correct bins. scripts/phase9_bin_accuracy_bench.py — the dedicated test research/LoRaTrace-Phases-7-10-Design.md's hardware matrix specifies ("Sweep calibration... known signals at low/mid/high bins"), separate from Endurance. Real sourced candidates at the low/ mid/high ends of the US region (LONG_SLOW 905.3125MHz, LONG_MODERATE 912.8125MHz, SHORT_SLOW 920.625MHz), each fired repeatedly (0ms-delay ARM, ~10Hz) through a live Sweep's whole duration to beat the ~40ms-per-bin dwell window: 11/24 attempts landed a real elevated BENCH_SWEEP_FLOOR reading, and all 11 were at the exact pre-computed bin index for that frequency — never a neighbor, never wrong. Sub-100% hit rate is expected (same dwell-vs- pulse-timing reality the 923MHz-edge work established), not a correctness gap; what mattered was 11/11 correct-bin attribution.

    • Endurance soak, scoped to 8 hours — found and fixed a real crash bug. No cited technical derivation exists anywhere in this project's docs for 24 hours specifically — it's a round-number target in the original design table (research/LoRaTrace-Phases-7-10-Design.md's hardware matrix), and 8 hours of back-to-back laps is already several thousand cycles, well past where a real leak or stability bug would be expected to surface. Documented as a deliberate, reasoned deviation, same convention Phase 7's own soak criterion was relaxed under once (docs/history/PROGRESS.md). scripts/phase9_soak.py, production firmware, WiFi off then on partway through (matching the design table's own "Off, then On" row).

      First 8-hour run: 4,143/4,148 laps completed, 5 failures. This was initially written up here as "0.1%, the project's own already- documented native-USB dropped-response pattern, not a device fault" — that was wrong, caught only by pulling session.csv off the SD card afterward and cross-referencing run-directory boundaries against the soak's own timeline. All 5 "failures" were actually identical, 100%-reproducible hard crashes:

      Guru Meditation Error: Core 0 panic'ed (Unhandled debug exception).
      Debug exception reason: Stack canary watchpoint triggered (logger)
      

      same backtrace every time, logger_task's own stack watermark plunging from 952B free at boot to 84B free before the first one. Each crash triggered a real RTC_SW_CPU_RST (software reset) and silently rebooted the device — which also fully explains the previously "unresolved" WiFi anomaly from the same run: WiFi wasn't buggy, the whole device rebooted and came back up in its normal boot-default (off) state, no code-path mystery required. Root cause: logger_task's 5,120-byte stack (logger_task.cpp, sized by inspection when added, never load-tested until this soak) was genuinely undersized for writeSessionRow()'s frame depth (a ~50-field SessionStats struct + a 320-byte row buffer, calling into SD/FatFS from the bottom of it).

      Fix: bumped to 8,192 bytes, matching wifi_task's own stack (a comparably deep SD/network call path) rather than guessing at another inspection-based number.

      3-hour verification re-run, same day: 1,507 laps, only 1 failure (a genuine isolated dropped response this time, not a crash — the device answered normally again the very next poll), zero Guru Meditation/RTC_SW_CPU_RST events, and WiFi stayed on continuously for 1,007 straight laps after being switched on with no reversion. Covers the timing of the first two original crashes (which hit at 0.34h and 2.81h into the original run) with margin. Not a full 8-hour re-confirmation, but real, clean, contradicting evidence against the bug recurring.

      Second full 8-hour run, with the logger fix in place: 4,353/4,353 laps, 0 failures, 0.0%. Zero Guru Meditation/RTC_SW_CPU_RST events anywhere in the log. One continuous run directory on the SD card (run0143) spanning the entire 8 hours confirms no reboot of any kind occurred, not just no crash. WiFi, switched on at the 4-hour mark, stayed on for all 2,340 remaining laps with zero reversions.

      "Bounded memory" directly confirmed from run0143's real session.csv: heap_free/heap_largest/heap_allocated_blocks all show a clean step function, not a decline — flat for the ~3.7 hours before WiFi turned on, one legitimate ~56KB one-time drop exactly when the AP started (its real allocation cost, matching Phase 7's own "recovered transient allocation, not a leak" distinction), then flat again for the remaining 4+ hours with WiFi running continuously. No drift in either phase.

      Timing tail, fully explained, not a bug. The recurring 19-35s laps (147/4,143 in the first run, a similar count in the second) are 100% explained: every single wp=0 (no peak found) lap took ≤3.5s; every single wp>0 lap took ≥4.2s, scaling with peak count (wp=2 laps cluster at 15-35s). This is Pass B correctly doing its documented job — up to 10 SF/BW combos' worth of CAD, plus a bounded 2.5s receive-on-hit window per combo that detects, for every real peak Sweep finds — not an anomaly. About 5% of sweeps in this room found something worth investigating; those sweeps take proportionally longer by design.

      radio_task stack margin, found and fixed the same way: pulling run0143's real numbers (not just logger's) showed radio_stack_free settling at a lifetime-minimum of 820B free out of 4,096 allocated (20.0%) within the first two hours and holding flat there for the rest of the run — not a leak, but below this project's own margin rule (25% or 1KB, whichever is larger). radio_task is the single most critical task in the system (owns the SX1262, must never block), so this was bumped proactively (4,096 → 6,144, proportionate to logger's own fix) rather than left at a margin already under the house rule just because it hadn't overflowed yet. A focused 2-hour verification run afterward — chosen to cover the ~1.9h mark where the old watermark hit its floor, with margin — came back 1,011/1,011 laps, 0 failures, 0 crashes, WiFi stable throughout.

      All five Phase 9 exit criteria are now closed, 2026-09-03.

Phase 10 (Field Analyzer) is accepted as planned scope. Whether it's required before v1.0.x was an explicit decision deferred until Phase 9 hardware evidence exists (ROADMAP.md) — that evidence exists now (above), and the decision was made 2026-09-03: Phase 10 is required for v1.0.x. Work is starting under an interim v0.10.x line, the same convention Phase 8/9 used while in progress; see ROADMAP.md's Phase 10 entry and Versioning table.

Phase 10 (Field Analyzer) — v0.10.1, all five exit criteria closed 2026-09-03:

  • Data layer, SCOPE_ACQUIRE, and the on-device UI (Meter/Waterfall/ Scope/Captures/Nodes) are done and hardware-verified, including a second hub — Tools, gating Probe/Sweep/Cell — added at the operator's request the same session (real scope beyond ROADMAP.md's own Phase 10 text, not a deviation from it). Two real bugs were found and fixed on real hardware during this pass: a whole-WaterfallHistory snapshot (~5.5KB) overflowing ui_task's 4096B stack (crashed the device opening Analyze > Waterfall), and the carousel-position footer reading straight off the raw UiPage enum ordinal instead of the main-carousel-relative index (would have shown "14/14" instead of "2/6"/"3/6" for the two new hub cards). Full writeup: ROADMAP.md's Phase 10 entry.
  • Memory budget — closed, confirmed two ways. ANALYZER_STATIC_BYTES (compile-time sizeof(), analyzer_budget.h) measures the four analyzer structures at 6,728 of the 8,192-byte incremental ceiling (82.1%, 1,464B headroom). A real 29-minute session.csv (run0006, below) confirms it end-to-end on hardware, not just at compile time.
  • Worst-case UI/radio run — closed. WiFi on, Sweep repeat mode running continuously, Waterfall open, for a full 60 real minutes. A background Serial Control watch (231 STATUS polls at 15s intervals) recorded zero dropped/unanswered requests and zero task_wdt/Guru Meditation signatures for the entire hour.
  • Outdoor and minimum-brightness readability — closed (operator check, 2026-09-03): confirmed good both in direct window sunlight and indoors.
  • A real bug the worst-case run itself surfaced — found and fixed same day, v0.10.1. Pass A found 50 energy peaks over that hour (STATUS's PBA=50), yet Waterfall showed nothing the whole time. Root cause: analyzerNoteSweepComplete() (analyzer_state.cpp, Core 0) read radio_task.cpp's live per-sweep peak-bin mask, but in repeat mode radio_task's own do-while loop calls straight back into performEnergySweep() for the next lap with no delay, and that lap's first line resets the same mask — Core 0's ~100ms poll cadence almost always lost that race, so every repeat-mode Waterfall row read an already-cleared mask regardless of what Pass A actually found. energy.csv itself was never affected (a separate, queue-based path that logs each peak the instant Pass A finds it). Fix: a second, stable snapshot buffer (energyPeakBinMaskAtComplete) taken atomically at sweep completion, read via a new radioEnergyPeakBinSetAtLastComplete() accessor; the Sweep page's own live occupancy ticks (drawSweepOccupancy()) are untouched, still reading the live mask on purpose so they keep updating progressively during an active sweep. Hardware-confirmed same day: operator re-ran Waterfall during a live repeat Sweep against real MeshCore traffic and confirmed hits now appear on the display.

Post-closure UI polish, same day (v0.10.2/v0.10.3), hardware- confirmed — real scope beyond the five exit criteria above, not a deviation: Waterfall gained a frequency axis (later merged into the plot box's own bottom border to remove a redundant line) and an Enter key that starts/stops repeat Sweep straight from the page. Meter gained a real bar gauge, an SNR line, and a right-column SF/BW/CR block — all real CaptureSummary data this page had access to and never showed — plus a range widened -30 -> 0dBm after a real -16dBm reading clipped flat against the original ceiling. Full writeup: docs/ROADMAP.md's Phase 10 entry.

Hardware finding, 2026-09-03 — found, isolated, and resolved by an SD card swap. While attempting the Stage 4 hardware verification above, every cold boot hit a 100%-reproducible task_wdt abort ~13-14s after [config] Applied channel override(s) — logger (Core 0) ran long enough during SD bring-up to starve IDLE0 past the watchdog's 5s window, hard-resetting the device into a boot loop, with [W] sd_diskio.cpp:180 sdCommand(): crc error firing immediately before it every time. Isolated by flashing the last tagged release (v0.9.0/b84d88c, no Phase 10 code) to the same device: identical crash, timing, and SD CRC warning — confirming this was the SD card, not firmware, in either version. Operator swapped the card same day; the same v0.10.0 build now boots cleanly — GPS clock sync, radio task start, and a Serial Control HELLO/STATUS round-trip all confirmed (STATUS reports SD=1), zero task_wdt/Guru Meditation signatures across a reset and a subsequent 20-minute passive Serial Control watch (40 STATUS polls at 30s intervals, 16:55-17:15 UTC, every poll answered, SD=1 throughout, no reboot). Confirms the fix holds under a sustained idle run, not just a single clean boot.

session.csv pulled off the card afterward (operator reattached it directly, 2026-09-03) — real confirmation, not just the compile-time number. run0006 covers a real 29-minute boot (uptime_s 3 through 1744, boot row through 29 periodic rows at the correct 60s cadence): analyzer_static_bytes,6728 on every single row, sd=ok throughout, zero row_drop/queue_drop/bus_miss/crc_err for the entire run, and heap_free/heap_min settling from 258,488B at boot to ~211,552B/ 207,108B within the first ~10 minutes and holding perfectly flat for the remaining ~19 — the same "one legitimate one-time settle, then flat" signature Phase 9's own soak established as healthy, not a leak. This closes the memory/telemetry side of Stage 4 for real, not just on paper.

Phase 11 (Cell) — added out of sequence, PARTIALLY hardware-verified: a bounded RSSI-only presence sweep of 869-894MHz (North American Cellular downlink), operator-requested after real wardriving runs picked up energy in that band near cell towers. It is not a decode of any kind — the SX1262 cannot demodulate GSM/CDMA/LTE — and not a fifth mission profile; same operator surface as Probe/Sweep (global hotkey C, dedicated carousel card), not a menu row. See docs/ROADMAP.md's Phase 11 entry and docs/DESIGN.md §5a for the full design and why it's numbered outside the normal phase sequence. Code and host-native tests landed 2026-09-01; on 2026-09-01 it was flashed to real hardware (v0.8.6, b7c845a) and confirmed: the C key runs a one-shot Cell scan end-to-end (toast through Cell: DONE with a real MHz/dBm reading on the carousel card, home channel restored after), and the Probe/Sweep mutual-exclusion guard holds in both directions (Cell during an active Sweep gives Sweep priority and refuses with Cell: UNAVAILABLE; Sweep during an active Cell scan is refused the same way). Still unverified: a real sweep near a known tower showing RSSI rising above the floor, and cell.csv/session.csv's new columns written correctly to SD.

Repeat mode (R), v0.8.7, hardware-verified 2026-09-01: Cell gained a repeat mode identical in shape to Sweep's own (radioRequestCellSweepRepeat(), back-to-back laps with a lap counter on the card). This also changed Sweep's own R key: it was a global hotkey (fired from anywhere, including with the menu open); it's now page-gated by ui_task.cpp to the Sweep/Cell cards specifically, matching how Enter (SELECT) already dispatches per-page — a no-op on any other page or with the menu open. Probe deliberately has no repeat mode (operator decision: "Repeat only on the Sweeps"). Confirmed on real hardware (b7c845a, same session as Cell's own verification above): Sweep-repeat still starts/stops correctly after the page-gating change; Cell-repeat starts/stops correctly with its lap counter rendering cleanly; R is a no-op on Probe (including mid-scan), every other page, and with the menu open from either Sweep or Cell; and mutual exclusion holds in both directions across a real repeat chain, not just a single shot (Cell refused during Sweep-repeat, Sweep refused during Cell-repeat).

FCC A/B block markers, v0.8.8, hardware-verified 2026-09-01: the Cell frequency bar now labels the FCC's own downlink sub-band split within 869-894MHz — Block A (869-880MHz + 890-891.5MHz) and Block B (880-890MHz + 891.5-894MHz), cited to 47 CFR § 22.905 (cell_plan.h's CELL_BAND_BLOCKS) — as a two-shade tick row under the bar with letter labels on the two segments wide enough to hold one. Regulatory block letter only, deliberately no carrier name (current licensee varies by market and isn't a fixed national fact — see docs/DESIGN.md §5a). Confirmed on real hardware (b7c845a): the tick row renders cleanly with no overlap against the lo/hi labels above or the disclaimer line below.

Workstream 12 (Focus Survey) — two gates closed as decisions

2026-09-06. W12's remaining gates were measurement-complete but decision-open, and they are now settled rather than left reading as unfinished measurement. Both follow one principle: Focus is a deliberate use, not a runtime state.

  • Radio-away budget: refused, conditionally. The exchange rate is measured and linear (predicted 0.459 against a measured 0.463), one Enter is one pass, and Activity's AWAY T card already shows the cost — so an operator self-governs against a number on screen. A cap would add a refusal path and a menu control for a runaway that cannot presently happen. The decision expires if Focus ever gains automatic repeat; that condition is recorded in focus_plan.h beside the code that would implement it.
  • Activity indication: refused. Focus reports coverage, never activity. A rule over the qualifying count works (C6 >= 2, 57/60, 95% CI [0.863, 0.983]), but its accuracy depends on link quality the device cannot know, and it measured worse where the band was busiest — failing toward false confidence exactly where an operator most wants it. The count and the full count-above-median ladder are still recorded, so a host with ground truth can conclude what the device may not.

The coverage campaign ran the same day (evidence): 80 passes, 80 completed with home restore, zero timeouts. Repeated passes are deterministic — observation-time stdev 0.0 ms and sample count invariant across every repeat at every dwell — so a coverage threshold in accumulated time and one in valid passes are the same statement, and the choice between them is presentational. Per-pass overhead is a flat 74 ms, confirmed independently of the §6.3 fixture, which is what argues for a minimum dwell rather than only a minimum accumulated time: 10 s of observation costs 13.0 s of Watch at 250 ms passes against 10.4 s at 2000 ms.

Coverage thresholds selected 2026-09-07 (src/focus_coverage.h): sampled = 1 valid pass / 2000 ms, repeated = 3 / 6000 ms, with a 500 ms dwell floor below which a pass counts toward neither. With the shipped 2 s pass that makes one Enter sampled and three repeated. sampled = 1 because Focus is one deliberate look; repeated = 3 is judgement and recorded as such, since the campaign showed every threshold equally achievable. The dwell floor is the part from measurement. focus.csv's coverage column is populated for the first time, alongside the raw counts it derives from — §3 forbids replacing those with a single word. Ten host tests cover it, most of them about what coverage must refuse to say: no label from an invalid pass, none carried across a retune, and none earned by repeating passes too short to mean anything.

The WiFi-off/on resource matrix is done — run0089's soak covered both conditions. What remains for W12: Portland field validation, release notes, and any companion-schema update. LOG_GUIDE.md now documents focus.csv, unblocked by the operator control the card-view work shipped.

8-hour soak (run0089) — memory bounded; three findings

2026-09-07, production 7a0b70c, 8.00 h of back-to-back sweeps with WiFi off then on, plus an accidental 7.6 h idle tail on battery (full writeup).

The gate passed. 4,112 laps, 0 failures, home restored every lap, every drop and error counter zero. Free heap moved in two discrete steps — UI canvas at 0.02 h, WiFi AP at 4.02 h — then sat at 153,928 bytes unchanged for 10.7 hours. heap_largest flat alongside it. The idle tail is what makes this conclusive: a leak would have shown where nothing was allocating.

Three things it caught that were not what it was looking for:

  • Pass-B spent 3.90 h of the 8 and promoted nothing (PBA=15680, PBD=0). Sweep away time is 80.2% Pass-B, 19.8% Pass-A; a 0-peak lap costs 833 ms and each peak adds 6–11 s. At 60.8% radio-away, the proportionality result from phase12 §6.3 implies ~2,650 packets not heard, for zero promotions. Not a defect — Pass-B promotes off-grid packets and the traffic here was on the home channel — but it is the cost half of a cost/benefit question, measured. Belongs to Workstream 17.
  • The UI task was at 93% of its 4 KB stack (ui_stack_free 288 B). -fstack-usage traced it to snapshot structs held as stack locals in draw functions (CaptureHistory 516 B, NodeRoster 672 B, and others). All are now function-static — only ui_task draws — taking the worst draw chain from 1,312 B to 592 B at a cost of 1.5% RAM. The stack itself went 4096 -> 5120 (60% used at the estimated peak, against 93% measured before either change); the old size had no recorded rationale, and now does. The 288 B was the deepest path exercised; pages not displayed during the run never contributed, so the true margin was thinner than measured.
  • The redraw instrumentation added the day before overflowed at 7.4 h. uiRedrawTotalUs was a uint32 of microseconds; the mean went visibly wrong at 8.01 h. Now uint64, with frame count exposed for windowed means. The max was unaffected: 223 ms worst frame, flat after 4.92 h.

Confirmed on hardware the same day (run0094, build 58410cb): with sweeps driven continuously and every card, every view, the Captures modal and the menu at full depth walked deliberately, ui_stack_free reads 1,624 B of 5,120 — 68% used, against 288 B of 4,096 (93%) before. Headroom improved 5.6x. The watermark settled within the first 65 seconds and never moved during the walk, so no card view is deeper than boot — which was the specific unmeasured worry. Redraw telemetry holds steady at ~60 ms mean with no wrap. Two honest corrections: the static fix moved the measured peak by 312 B rather than the 720 B predicted from summing frames, so the drawPage+drawMeterPage chain was never the deepest path; and -fstack-usage on this project's own files cannot see the frames that make up the rest of it.

What's still open

  • v1.1.0 audit (2026-09-07): the report is here; its repair status is the dated section at the end of it. The first repair pass has landed in the working tree and is verified only by host tests and a clean build — no hardware run yet. Cell's missing per-bin retune (A01) is fixed in code; Cell results from released v1.1.0 remain unreliable and any Cell data captured on that build should be discarded, not reinterpreted. A second pass the same day closed A10–A13, A15, A17–A19, A24 and A25, with the AP credential, CSRF defenses, downloads, run manifest and session id exercised on hardware over the device's own AP. A third pass closed A16, A22, A26 and A27, and the repair half of A21. Fuzzing (A27) found two real defects in the safe-export written the same day. Still open: key-import gates (A14, nothing to gate yet), A21's bounded command arbiter (Workstream 13 scope), and per-action budgets for Probe and Sweep, deliberately not set without measurement. A01 is fixed and confirmed on hardware (run0117/run0119): a 5.6 s lap matching the documented full-begin() timing, 40 dB of structured spread across bins, no radio errors. A23 is fixed and confirmed (run0119): 404 of 404 Cell rows written, drops 0, against 192 of 404 lost before.

  • A29 found and fixed the same day (2026-09-07). With tuning and row loss repaired, Cell's strong bins formed a 0.75 MHz comb whose phase moved randomly between laps. A controlled sweep on the bench image (two interleaved reps per value) showed Cell was sampling RSSI before the receiver had settled after each retune: at 0 ms only 33 of 101 bins produced a real reading, at 5 ms all 101 did, and 10 ms added nothing. Fixed by CELL_SETTLE_MS = 5 (~0.5 s per lap). Cell data captured before this fix is unusable, including the "quiet" readings — the settled median is −91 dBm against −117 dBm unsettled. Confirmed on the production image (run0127): 101 of 101 bins written, 99 above −110 dBm, median −95.3 dBm, no comb, 0 drops.

  • Bench SD card / boot-loop finding — resolved 2026-09-03. The bench Cardputer's task_wdt boot-loop (see Phase 10 section above) was the SD card, not firmware; operator swapped it same day, and a clean boot plus a 20-minute zero-crash passive Serial Control watch confirmed the fix. Bench device trusted again.

  • Phase 10 (Field Analyzer) hardware confirmation — closed 2026-09-03. All five exit criteria done, including a real bug (a repeat-mode Waterfall race condition) the worst-case run itself surfaced and got fixed the same day; see the Phase 10 section above for the full writeup.

  • Sweep silence near real MeshCore traffic — diagnosed, 2026-09-03: it's dwell timing, not the noise floor. Originally opened as "is 35.0dB too conservative?" after a worst-case run found 50 Pass-A peaks but zero packet promotions while Watch/Trace saw plenty of real traffic from a 6-foot-away MeshCore repeater (operator's own pyMC_Repeater node, real local mesh). Three independent lines of evidence closed this out:

    • The floor itself is fine. The repeater's own noise-floor monitor reports -94.0dBm average at this exact location; its packets average -34.8dBm — a ~59dB clearance over the floor, well past the 35dB margin. Real traffic here is nowhere near the threshold.
    • Watch/Trace has no reception bug. A live correlation (Serial Control's new RXP/RXC fields, radioPacketCount()/ radioCrcErrorCount(), added this session) logged 96 real packets in 6 minutes with zero CRC errors — more than the repeater's own log showed for the same window, since Trace hears the whole broadcast mesh, not just one node's vantage point. (An earlier reading of this session mistakenly used R, which is Probe's own recovery counter, not a packet count despite the letter — corrected before drawing any conclusion from it.)
    • Direct RTL-SDR ground truth confirms the dwell-timing theory. A focused 909-912MHz capture during 19 live single-shot Sweep laps showed near-continuous real bursts at 910.5MHz (roughly one every 1-8 seconds, 182 flagged events over ~5 minutes) — yet Sweep only registered a peak on 5/19 laps (26%), even with transmissions that dense nearby. A ~40ms dwell at one bin per lap only rarely coincides with an independently-timed burst; this is a receive-window problem, not a sensitivity problem. Matches (with much denser real traffic) the same effect the 923MHz-edge bench work already established for a single controlled signal.
    • Bonus, unplanned finding: the same test showed Trace decoding only 6 packets in ~5 minutes while Sweep laps ran back-to-back, versus 96 in 6 minutes with the radio otherwise free — a real, now-quantified ~15x drop in Watch/Trace's catch rate while Sweep (repeat or back-to-back single-shot) monopolizes the radio. Not a bug, an inherent trade-off worth knowing about, not previously measured.

    No firmware change was indicated by this investigation itself — the margin, the reception path, and the dwell design were all working as built for the case tested. RXP/RXC stay on STATUS going forward — genuinely useful for future "is Trace actually receiving" questions, not just this one.

    • Margin operator-configurability — done, v1.0.1 (2026-09-03, same day, a separate follow-up report): a field test at real deck range (-55 to -72dBm) has far less clearance over the floor than this investigation's 6ft/59dB test did, so the 35dB default can plausibly still gate a weak, distant, clean-SNR signal even though it wasn't the cause here. System > Tuning > Margin, 15.0-50.0dB/5.0dB steps, /loratrace/sweep_margin.txt.
    • Per-bin retune cost — cut ~4.1x, v1.0.2 (2026-09-03/04, same session): performEnergySweep() was calling a full radio.begin() (RadioLib -> modSetup(): hardware reset, chip re-detect, TCXO restart, full config reload) for every bin, even though SF/BW/CR/sync never change across a sweep — only frequency does, and RadioLib's own setFrequency() only re-runs image calibration past a 20MHz jump (this sweep's 250kHz step never crosses that). Replaced with one real begin() per sweep (and again right after any bin that ran a Pass-B CAD attempt, which leaves the radio on that candidate's own SF/BW/ sync) plus standby()/setFrequency()/startReceive() — three single-SPI-command calls — for every other bin. Hardware-confirmed at a fixed 35.0dB margin (zero Pass-A peaks either side, so zero Pass-B time muddying the comparison): four back-to-back 85-bin US-region sweeps before vs. after — 3471/3435/3475/3472ms (avg 3463ms, ~40.7ms/bin) vs. 834/866/833/866ms (avg 850ms, ~10.0ms/bin). The earlier, less-isolated attempt at this same measurement (at a then-sensitive -20dB margin) is why it needed isolating: several Pass-A peaks per sweep triggered Pass-B CAD, and that inflated results up to 14-20s — which is the next finding, below.
    • New, bigger cost found while measuring the above: Pass-B's own bounded receive-on-hit window. passBCadAtBin() reuses Probe's CAD
      • bounded-receive-on-hit sequence at each of the first PASS_B_MAX_PEAKS_PER_SWEEP (8) Pass-A peaks per sweep; a CAD hit opens DISCOVERY_RX_WINDOW_MS (2.5s) waiting for a full packet. At a sensitive margin finding several real peaks, that's up to ~20s added on top of the now-~850ms Pass-A scan — an order of magnitude bigger than the per-bin cost just fixed above. Evaluated same session (operator call: "I think that's fine") and left as-is — not fixed, tracked here as the next lever if sweep duration at a sensitive margin becomes a real problem. Dwell timing itself (the RTL-SDR- confirmed root cause of the original "Sweep silence" case above) is still open as a design limitation, not a bug — the freed per-bin time hasn't yet been reinvested into more samples/bin; that sizing was deliberately deferred to a real before/after measurement rather than guessed (see radioEnergyLastAwayMs(), already exposed on STATUS's EA field and the Sweep card, for exactly that measurement).
    • Reinvesting the freed time into more samples/bin — sized and shipped as ENERGY_SWEEP_SAMPLES_PER_BIN = 34 (radio_task.cpp), same session: an isolated N=1 hardware measurement put fixed per-bin overhead at ~6.98ms/bin, each additional 1ms-spaced sample at ~1.0ms/ bin, and 34 lands total per-bin cost back at ~39.9ms/bin — the same total sweep duration as before the retune fix, just ~34ms of it now real dwell instead of ~3-4ms. CELL_SAMPLES_PER_BIN (4) keeps performCellSweep() — never touched by this retune work — at its original dwell, since sharing the constant would have silently stretched Cell's own sweep by the same ~30ms/bin as a side effect.
    • Real-traffic correlation against pyMC_Repeater: inconclusive between wide-dwell and narrow-dwell, not a confirmation of either. Two back-to-back ~4.1-minute windows, each driving bounded sweeps single-shot from a host script (radioEnergyPeakCount() returns the live counter, not a stable per-lap snapshot — true on-device repeat mode's own no-delay next-lap loop can reset it before a poll reads the terminal value, so polling only between host-orchestrated bounded sweeps sidesteps that race) against the operator's pyMC repeater's live /api/stats as independent ground truth (docs/hardware-results/private/dwell-reinvest-20260904T005831Z-wide- dwell-N34.results.jsonl and ...T011142Z-narrow-dwell-N4.results.jsonl, gitignored — real packet content):
      • Wide dwell (shipped ENERGY_SWEEP_SAMPLES_PER_BIN=34, ~3.4s/lap quiet): 28 laps, 8 whole-band peak events against 35 real packets in-window = 22.9% (95% CI 8.9-36.8%).
      • Narrow dwell (N=4, ~0.85s/lap quiet): 63 laps — 2.25x more revisits in the same wall-clock window — 11 whole-band peak events against 44 real packets in-window = 25.0% (95% CI 12.2-37.8%).
      • The two intervals overlap almost entirely — this pair of runs cannot distinguish the conditions. An earlier read of the wide-dwell run alone (before the narrow-dwell follow-up existed) reasoned from real packet airtime (142-490ms, median 244ms — one to two orders of magnitude longer than either dwell width) that revisit frequency should dominate over dwell width, and predicted narrow-dwell's 2.25x more revisits would show a materially higher catch fraction. It didn't, at least not distinguishably in this sample — that airtime-based theory is not confirmed by this data, and neither is the original "wider dwell helps" premise the whole reinvestment was built on. Both conditions also land in the same range as the original RTL-SDR-confirmed ~26% figure at the old pre-any-of-this- session dwell.
      • Real limitations on trusting this too far either way: WP is a whole-band count with no per-bin attribution over Serial Control (a "peak" isn't confirmed as pyMC's own 910.525MHz traffic specifically, just the closest available proxy — the same gap that made the first attempt at the plain retune-timing measurement need isolating from Pass-B noise); n=8 and n=11 hit events are both too small to resolve anything but a large effect; and the two windows sampled different moments of real, non-stationary ambient traffic (44 vs. 35 packets), not a repeatable controlled source. A conclusive answer would need either many more repeat trials or bin-specific ground truth (RTL-SDR, the same tool the original dwell-timing root-cause finding used and this follow-up deliberately didn't reach for).
    • Resolved, v1.0.3 — it was the wrong variable all along; capture is a timeshare problem, not a dwell problem. Asked to get capture to a usable rate, the model that settles it is simple: decoding a packet needs the receiver parked on its frequency when the preamble lands AND held there for the packet's whole airtime (142-490ms measured, median 244ms). No full-band sweep can offer any single bin that, at any dwell width — 85 bins means each gets a few ms per lap. So capture is bounded by the share of wall-clock time parked on the channel, which is why ENERGY_SWEEP_SAMPLES_PER_BIN could never reach it and why both arms of the A/B above landed in the same range. Measured directly, with repeat Sweep running continuously and pyMC as ground truth (docs/hardware-results/private/capture-rate-*): 0 of 42 real packets captured, 0.0% — not the ~15x degradation recorded earlier but effectively total blindness while sweeping. The fix is to timeshare the radio rather than tune it: performEnergySweepHomeListen() parks on the home channel with RX armed for the capture window (2000ms default) after each lap in repeat mode, servicing packets through HOME_LISTEN's own readDetectionLocked()/enqueueDetection() path — same Detection, same RXP, same detections.csv row, no second definition of "a real packet". Single-shot Sweep is untouched. Immediately-following window, same conditions: 22 of 27 captured, 81.5%, zero CRC errors, lap time unchanged at ~833ms, full-band survey cadence ~2.9s instead of ~0.9s. ENERGY_SWEEP_SAMPLES_PER_BIN reverted 34 → 4 in the same change, the reinvestment having never been justified by real traffic. Shipped with an on-device toggle in v1.0.4 (System > Tuning > Capture: Off/1s/2s/4s, persisted to /loratrace/capture.txt), since the cadence-vs-capture trade is an operator call, not a constant. Pooling both treatment runs gives the honest figure: 30/44 = 68% (95% CI 54-82%) — the single-run 81.5% was optimistic at n=27. v1.0.5 closes the display half of the same gap: Waterfall marks those captures green (energy stays yellow), so the page can no longer read "QUIET" while packets are being decoded — the visible symptom this whole investigation started from. Verified over 45 consecutive rows (bin 34 = 910.525MHz, 15 rows carrying 21 packets).
  • Cell hardware verification (Phase 11, above) — C key, mutual exclusion against Probe/Sweep (both directions), and the carousel card are now confirmed on real hardware (2026-09-01). Still open: confirm a real cell-band RSSI reading actually rises near a known tower, and confirm cell.csv/session.csv's new columns write correctly to SD.

  • Phase 9's endurance soak — closed 2026-09-03. All five exit criteria are done; see the Phase 9 section above for the full soak writeup (a real crash bug found and fixed along the way).

  • Pass B's other eight SF/BW combos still have only n=20/condition (UNVERIFIED) — more bench cycles would be needed before extending STRONG/NOISY past the two combos that have it now (docs/research/phase9-sweep-pass-b-design.md's open questions).

docs/history/PROGRESS.md has a much older "Open questions" list dating back to Phases 1-2 (2026-08-22 through 2026-08-27). Most of those are resolved and marked [x] there; a few unresolved ones (MeshCore's encryption/PSK scheme, CAD symNum tuning, Meshtastic's exact per-slot US frequency table) were never revisited after that file stopped being the live status doc — treat them as leads to re-check, not confirmed-current open items.

Three hard-won hardware facts

Kept here as well as in CLAUDE.md (agents read that file, humans are more likely to land here first):

  • The IO expander's P0 powers the GPS as well as switching the RF antenna path. A "dead" GPS or a silent radio is often just this.
  • A wrong sync word is silent, not loud — the radio simply never interrupts, while still hearing unrelated traffic that matches.
  • A same-node-id detection pair with wildly different RSSI seconds apart is very likely a genuine mesh relay, not a logging bug — packet_id/hop_limit/hop_start/relay_node in detections.csv prove it either way.