Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
21 commits
Select commit Hold shift + click to select a range
910bc3a
Archive the completed v1 roadmap, open the V2 gate board
d3mocide Sep 5, 2026
ded5720
Workstream 12: bounded Focus Survey, bench evidence, controlled matrix
d3mocide Sep 5, 2026
2c2577f
Replace Phase 12's stale implementation order with ordered next steps
d3mocide Sep 5, 2026
781221c
UI Mockups
d3mocide Sep 5, 2026
c699fb0
Raise the Focus bench sample ceiling; warn off /dev/ttyACM* targeting
d3mocide Sep 5, 2026
d0a9ec4
Pin the bench home channel into the image; harden the bench harnesses
d3mocide Sep 5, 2026
f4572ac
Derive Focus's sample count from the dwell, at a measured 20 ms spacing
d3mocide Sep 5, 2026
0f7bc7a
Add a WiFi control bridge to the bench transmitter
d3mocide Sep 5, 2026
fa79281
Measure Focus's qualifying RSSI condition at field levels, and reject it
d3mocide Sep 5, 2026
0138a54
Measure what Focus costs Watch, and find the cost is exactly its away…
d3mocide Sep 5, 2026
8128be7
Count samples above an adaptive floor; it works where statistics did not
d3mocide Sep 5, 2026
a4bc335
Bound the count rule: a fraction of samples, and occupancy is the limit
d3mocide Sep 5, 2026
b74b4b5
Reconcile STATUS and CHANGELOG with the dwell and occupancy sweeps
d3mocide Sep 5, 2026
9d5a286
Second baseline withdraws the occupancy conclusion
d3mocide Sep 5, 2026
7237186
Reconcile STATUS and CHANGELOG with the withdrawn occupancy conclusion
d3mocide Sep 5, 2026
9be6b65
Log a Sweep/Waterfall sampling review as a candidate workstream
d3mocide Sep 5, 2026
b749ae7
A single packet is detectable; reconcile the design doc, three commit…
d3mocide Sep 5, 2026
2a3a352
Workstream 17 entry: Pass A reads a floor ~14 dB low as shipped
d3mocide Sep 5, 2026
e8e3dd4
Pass A's floor stabilises by 3 ms, at about 15% of lap time
d3mocide Sep 5, 2026
44a4c94
Pass A misses real repeater traffic because of the settle under-read
d3mocide Sep 5, 2026
c16a7b5
Confirm at 300 laps: shipped Pass A flags neither busy channel
d3mocide Sep 6, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .github/workflows/build.yml
Original file line number Diff line number Diff line change
Expand Up @@ -78,7 +78,7 @@ jobs:
# Rolling "latest main" build, separate from the versioned vX.Y.Z
# releases in release.yml. Overwrites the same dev-latest tag/release
# on every merge so there's always a one-click, no-login download for
# SD-drop/Launcher testing — see docs/ROADMAP.md Distribution section.
# SD-drop/Launcher testing — see the root README's install guidance.
# Only on pushes to main (not PRs, not every commit on a branch).
- name: Move dev-latest tag
if: github.ref == 'refs/heads/main' && github.event_name == 'push'
Expand Down
20 changes: 11 additions & 9 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,19 +14,21 @@ should be fixed to match, not the other way around.
Also relevant, referenced from CLAUDE.md:
- [docs/STATUS.md](docs/STATUS.md) — current status, what's
hardware-verified, and what's still open.
- [docs/DESIGN.md](docs/DESIGN.md) — the "why" behind RF/architecture
decisions.
- [docs/ROADMAP.md](docs/ROADMAP.md) — phase-by-phase build order and
scope.
- [docs/DESIGN.md](docs/DESIGN.md) — the "why" behind the shipped v1
RF/architecture foundation.
- [docs/ROADMAP.md](docs/ROADMAP.md) — active V2 workstreams, gates, and
release policy.
- [docs/research/V2_DESIGN.md](docs/research/V2_DESIGN.md) — V2 product
boundaries and design direction; read it before V2 architecture work.
- [CHANGELOG.md](CHANGELOG.md) — short, ongoing changelog; full
pre-2026-08-29 session-by-session decisions log is archived at
[docs/history/CHANGELOG.md](docs/history/CHANGELOG.md).
- [SECURITY.md](SECURITY.md) — known attack surface and how to report issues.
- [docs/HARDWARE_TESTING.md](docs/HARDWARE_TESTING.md) — repeatable bench
matrix and Phase 7 memory acceptance rules.

**Don't read `docs/history/PROGRESS.md` or `docs/history/CHANGELOG.md`
end-to-end by default.** They're a frozen pre-2026-08-29 development log,
not required context for every task — `docs/STATUS.md` already gives the
current-state summary. Search them for the specific date/version/topic you
need instead of reading front to back.
**Don't read `docs/history/PROGRESS.md`, `docs/history/CHANGELOG.md`, or
`docs/history/ROADMAP_V1.md` end-to-end by default.** They're historical
records, not required context for active V2 work — `docs/STATUS.md` and the
active roadmap give the current-state summary. Search history for the specific
date/version/topic you need instead of reading front to back.
224 changes: 224 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,202 @@ project (not a log of how it got there), see [docs/STATUS.md](docs/STATUS.md).

## 2026-09-05

- Measured the positive case the withdrawal left open: **a single packet is
detectable**. One armed 148 ms transmission placed inside a 2,000 ms pass —
7.4% occupancy — was caught 29/30 with 0/30 false positives at the stronger
link. A 148 ms packet sampled every 20 ms yields about seven elevated
samples against the four the rule needs, so nothing about the instrument
prevented it; the first link did. Not unconditional, though: the same rule
at the weaker link failed on sources filling four times as much of the pass,
and the companion position carrying real traffic measured *worse* — a
-57 dBm event in a control trial both produced false positives and
suppressed counts by lifting the median, so Focus is least reliable exactly
where a band is busiest. Also discarded a placement analysis that tried to
time the pulse against the window from log timestamps: `TX_STARTED` reaches
the host only when the harness next polls, so it records host polling rather
than RF timing, and its output was self-contradictory.

- Ran a second baseline at ~25 dB SNR (600 trials) and **withdrew** the
occupancy conclusion drawn from the first. Replacing the transmitter's
stubby with a matched whip and standing both antennas vertical gained ~13 dB
— vertical alignment mattered more than the antenna, the horizontal pair
having sat in each other's pattern null at 918.5 MHz. At 28.6% occupancy
that took detection from 37% to 87%, so the "cliff" was a property of the
first link, not a statistical limit of the instrument. The claim that the
rule "detects a persistently occupied channel, not individual packets" is
therefore withdrawn, along with the argument that the evidence favours CAD
over sampling for §3's activity basis. Two conclusions did survive both
links: the 100 ms six-sample floor (13 dB bought almost nothing) and the
rule's fraction-of-samples form. Recorded as a method note too — three
conclusions came from one link, two held and the most confident one
inverted.

- Bounded the count rule with two more sweeps (480 trials, no arm failures).
It is a **fraction** of accepted samples (~4-5%), which transfers between
500 ms and 2,000 ms passes; no threshold rescues a 100 ms pass, whose six
samples cannot both catch the source and reject ambient — so a short pass
may report coverage and must not report activity. The occupancy sweep is the
consequential one: detection runs 90-93% while the source fills 43-57% of
the pass and collapses to 37-43% at 28.6%. A single SF8 packet inside a
2,000 ms pass is 3-7% occupancy, far below that. **The rule detects a
persistently occupied channel, not individual packets**, which is the honest
scope of an RSSI-sampling instrument and makes CAD or packet reception the
better-supported basis for §3's observed activity. Median-as-floor held at
every occupancy tested, but only because a weak source leaves most samples
reading like noise; a strong source would break it and that regime was not
reached.

- Found an activity basis that survives field levels. After every RSSI summary
statistic was rejected, the remaining idea was to count samples above an
adaptive floor rather than take an extreme of them — `peak` is one sample and
noise-prone; a count integrates. Bench images now report a ladder of counts
at median + 2/4/6/8/10/15/20 dB, so one run evaluates any such rule offline
instead of reflashing per candidate. Across 120 trials at the same porch
configuration, `C6 >= 2` detected 57/60 source-on (95% CI [0.863, 0.983])
against 1/60 source-off — and that one flagged control trial read -63 dBm
against a -101 dBm median, a real transmission that was not ours, so the
false rate is an upper bound rather than a measurement against silence.
Still a candidate, not a constant: the threshold is a count out of 101
samples, and since sample count now scales with dwell, "2 samples" is 2% at
2,000 ms and 33% at 100 ms. `qualifying_count` stays unpopulated until that
is expressed as a fraction or validated per dwell.

- Ran Workstream 12's §6.3 Watch-opportunity comparison. Two 240 s arms
against one independently timed reference train: Watch alone received 0.883,
Watch with Focus interleaved at a 48.1% away fraction received 0.463,
intervals non-overlapping, zero CRC errors either side. The loss is
proportional to away time and nothing else — 0.459 predicted from the
baseline against 0.463 measured — which means Focus's recorded away duration
is an honest proxy for what a request costs Watch, with no hidden retune or
recovery penalty. It also means there is no mitigation: half the listening
time away is half the packets, with restored-Watch gaps reaching 7.45 s. The
away-time budget decision stays open on purpose; the measurement gives the
exchange rate, not the policy.

- Measured Focus's qualifying RSSI condition at a realistic signal level and
**rejected it**. The `p90 >= -90 dBm` candidate came from a bench source ~70
dB above ambient, where every metric separates and the choice looks easy.
Repeated with the transmitter outdoors (source peak ~-85 dBm against a
-99/-100 dBm floor, 120 trials, zero drops), it fails — and so does every
floor-relative variant tested: one of ten metric/position combinations
separates, by 2 dB, which is inside ordinary RSSI variance. Two reasons: at
low SNR `p90` collapses even when the source radiates most of the window,
and real MeshOregon traffic reached -94 dBm during control trials, which no
RSSI condition can distinguish from a controlled source. Coverage reporting
is unaffected; what the evidence refuses is the step from "RSSI was
elevated" to "something transmitted", which §3 forbids anyway. Remaining
routes: a per-pass count above an adaptive floor (`qualifying_count` is
reserved and unpopulated), or CAD/packet evidence.

- Selected Focus's sampling policy from measurement instead of reasoning:
`FOCUS_SAMPLE_SPACING_MS = 20`, with the sample count derived from the dwell
(`ceil(dwell/spacing) + 1`) rather than fixed at 8. A 360-trial sweep held
dwell at 2,000 ms and varied only the spacing: a 94 ms source was missed at
286 ms and 100 ms spacing (8/15 and 2/15) and caught 15/15 at both 50 ms and
20 ms, with worst-case reading improving from -97 to -66 dBm between them.
Detection collapses once spacing approaches the source's airtime. Finer
sampling turned out to be free -- radio-away measured 2,073-2,075 ms whether
a pass took 8 samples or 101 -- so the only real argument for coarse
sampling was one nobody had checked. Evidence appended to
`docs/hardware-results/2026-09-04-phase12-focus-matrix.md`.

- Ran Workstream 12's §6.2 controlled dwell matrix: 900 trials in 54.7 min,
zero transport errors, zero drops, home restored on every trial; 14 of 15
arms separated a controlled source from ambient 30/30 vs 0/30. The finding
that matters is a limit — **detection tracks the source's airtime against
Focus's sample spacing, not dwell length**. With a fixed 8 samples, spacing
is dwell/7, so a 2000 ms dwell observes eight instants rather than 2000 ms;
worst-case detection degraded monotonically as source airtime approached
that spacing and one arm stopped separating. That is §3's "observation time
is not coverage" rule with numbers behind it, and it means a later slice
should scale samples with dwell. Bin-center offset turned out not to be the
dominant term after all. No single fixed RSSI condition separates every arm,
so §3.1 gets a candidate rather than a constant. Summary in
`docs/hardware-results/2026-09-04-phase12-focus-matrix.md`.

- Two methodology bugs found by running the §6.2 matrix rather than by
reading it. A single armed pulse at a fixed delay cannot cover every dwell
arm — at 100 ms the pulse began after the window closed, and source-on was
indistinguishable from source-off across 18 trials. Re-arming continuously
fixed that but at SF12's ~275 ms airtime the sends overlapped and their tail
bled into the next trial, making source-off read -26 dBm. The runner now
paces the burst by the transmitter's own `TX_DONE` and waits for the tail to
clear before a trial ends: 100 ms arms then separate cleanly (-25/-27 dBm
against -97/-101 dBm). Both failure modes are recorded in the design doc so
the next harness doesn't rediscover them.

- Added `scripts/phase12_watch_opportunity.py` for §6.3, the measurement that
gates Focus's maximum radio-away budget: Watch-only versus Watch-with-Focus
against one independently timed pulse train, reported as Wilson intervals
plus away time, completed requests, and restored-Watch gap distribution. It
refuses to run unless the device's home channel matches the reference train,
since otherwise both arms hear nothing and their equality would read as a
reassuring result, and it reports the away fraction alongside the loss
because a small loss at a small duty cycle predicts nothing about a larger
one.

- Built Workstream 12's §6.2 matrix tooling: `scripts/phase12_focus_matrix.py`
runs paired source-on/source-off Focus trials across dwell arms and writes
one durable JSONL row each, and `scripts/phase12_focus_matrix_report.py` does
the offline analysis — per-arm distributions, the lowest separating 1 dB
condition or an explicit "these overlap", and Wilson 95% intervals rather
than bare proportions. Neither can emit a coverage label. Validated on
hardware with a 4-trial smoke run at 912.750 MHz: source-on P90 -66/-68 dBm
against -97/-98 dBm quiet, 573 ms radio-away, no drops.

- Found that none of §6.1's three fixture frequencies sits on a Focus bin
center. Focus tunes at the home channel's bandwidth (125 kHz here), so the
offset changes what a pass can observe: low and mid are +62.5 kHz off bins 13
and 43, and the high position sits exactly on a bin boundary, 125 kHz from
either neighbour — outside the passband. A null there would measure geometry,
not sensitivity, so the matrix now also runs two exactly bin-centered
controls from the same transmitter table, and every trial records its offset.

- Closed Workstream 12's last Engineering-gate item by measuring the Focus
budgets off the built image instead of estimating them
(`docs/research/phase12-survey-truth-design.md` §4.2): 12 B request, 40 B
result, 148 B histogram, 188 B working state against a 256 B target, 159 B
of Focus statics, a 1,072 B `radioTask` frame in a 6,144 B stack, and a
189 B worst-case `focus.csv` row against a 256 B buffer that drops rather
than truncates. Everything landed inside the bounds §4.1 had set, so
nothing needed resizing. The row buffer now comes from one shared
`FOCUS_CSV_ROW_MAX` and a test formats the saturated row, so the writer and
its budget can't drift the way `STATUS`'s hand-sized buffer did.

- Phase 12: bounded a Focus request in wall-clock time, not just in samples.
Each sample waits on the shared SPI bus, so a contended bus could stretch a
pass past its dwell with the radio away from home and nothing to stop it —
`focusRuntimeTimeout()` existed but no code path reached it. A request now
carries a dwell-plus-slack deadline and terminates as `timeout`, restoring
home like every other exit. Bench-only additions make the remaining Device
gate reproducible: `BENCH_FOCUS`'s optional 4th field arms a one-shot
sample-loop stall (production only times out under real contention), and
`BENCH_ACTION` starts/cancels/reports Cell and Scope, which are menu-only in
production, so Focus's mutual exclusion against them can be driven by a
fixture. Verified on hardware the same day: a stalled request timed out,
restored home, and wrote its durable row inside a 1,573 ms radio-away
window, and Cell and Scope each arbitrated with Focus in both directions.

- Fixed a latent silent-drop in Serial Control's `STATUS`: its argument buffer
was hand-sized to 240 bytes against a frame budget of ~230, and
`serialControlFormatFrame()` drops an over-long frame rather than truncating
— so a long session's wider counters would have lost the whole status frame,
newest fields first, with no error. The argument budget is now derived from
the frame size, the frame limit is 384, and a host test formats at
saturation. The formatter also builds in place instead of staging through a
second full-size buffer, which nets less stack on the 4KB UI task than
before.

- Phase 12 engineering: wired one bounded, bench-only Focus Survey through
Core 1, the Core-0 `focus.csv` writer, and restore-before-publish terminal
state. Added framed `BENCH_FOCUS` and GPS-free compact
`BENCH_FOCUS_RESULT` readback, durable-write health counters, and the
explicit-transmit-gated `scripts/phase12_focus_bench.py` plus
non-transmitting `scripts/phase12_focus_behavior_bench.py` fixture
harnesses. Quiet/source-on, cancel, injected-failure, and Probe/Sweep
arbitration hardware checks completed with successful home restore and SD
rows; they do not set a coverage/activity threshold.

- `pages.yml`'s "Download stable release assets" step now retries (6x/20s)
instead of failing hard: a maintainer publishing a draft release before
release.yml's build job finishes uploading assets makes `release:
Expand Down Expand Up @@ -78,6 +274,34 @@ project (not a log of how it got there), see [docs/STATUS.md](docs/STATUS.md).

## 2026-09-04

- Archived the completed v1 roadmap at the immutable `v1.0.7` tag and
streamlined `docs/ROADMAP.md` into the active V2 gate board. Updated
`CLAUDE.md`, `AGENTS.md`, and documentation indexes so future work opens
V2 gates/design first and searches v1 history only when needed.

- Opened V2 Workstream 12 (Survey truth) at Design entry: one selected-bin
Focus Survey, fixed streaming statistics, durable `focus.csv` contract,
controlled transmitter/RTL-SDR truthfulness matrix, and Portland metro,
Oregon as the privacy-preserving field-validation area. No firmware scope
has entered implementation; radio-away and coverage thresholds remain
evidence-gated.

- Began Workstream 12 Engineering with host-tested, radio-free Focus request
and `focus.csv` contracts: a single sourced bin, 1 dB bounded RSSI histogram
(median/P90/peak), explicit completion states, and an intentionally blank
`coverage` column. No menu, radio task, logger, or coverage conclusion is
wired yet.

- Tightened the Focus bench-request contract before radio integration: exactly
one pass, 2--2,000 ms dwell, and 2--64 samples. Added a host-tested state
machine that records restoration before a terminal result and makes restore
failure override partial/cancelled work.

- Wired the bench-only Focus vertical slice: a Core-1 one-bin RSSI request,
mutually exclusive ownership/restoration, four-row Core-1-to-Core-0 queue,
durable `focus.csv`, and framed `BENCH_FOCUS bin:dwell_ms:samples` control.
The production image rejects the command; hardware evidence is still open.

- Ran a `/code-review` pass and then a whole-project audit
(`docs/research/2026-09-04-project-audit.md`), and fixed what it found
(`v1.0.6`). Radio-ownership bugs in v1.0.5's new capture window:
Expand Down
Loading
Loading