From 5407dd2eb677effcf6774d73c6ae01894c1d4d33 Mon Sep 17 00:00:00 2001 From: Claude Date: Sun, 4 Oct 2026 03:39:34 +0000 Subject: [PATCH] Record wave-4 status, audit notes and next-session handover Co-Authored-By: Claude Sonnet 5.5 Claude-Session: https://claude.ai/code/session_01VqsBNqvaWezf1DwQrAnVgX --- AUDIT_OPEN.md | 32 +++++++++++++++++++++ NEXT_SESSION.md | 35 +++++++++++++++++++---- docs/SHIP_RELEASE_PROGRAMME_2026-09-21.md | 18 ++++++++++++ 3 files changed, 79 insertions(+), 6 deletions(-) diff --git a/AUDIT_OPEN.md b/AUDIT_OPEN.md index 19709ec..05002b4 100644 --- a/AUDIT_OPEN.md +++ b/AUDIT_OPEN.md @@ -1,5 +1,37 @@ # Open audit items +## 2026-10-04 — wave 4 (S07, S13, S15/S17, S19, S21, S22, S24, S26, S27) + +All off by default or observe-only; flags are in the programme's wave-4 table. + +- **S27 Teploy adapter (PR #54):** flag off. When on, it is stricter than the + legacy path (unreadable status, including a brand-new destination, holds). + A manual `teploy deploy` landing inside Ship's deploy is overwritten and still + reads back confirmed; the lease serialises Ship's own releases on one host + only. Upstream feature requests (not filed): deploy dry-run; deploy generation + and compare-and-set in `status --json`; documented `logs` flags and a + no-deployment-versus-unreachable signal. +- **S21 (PR #51):** provenance is a sidecar, not fields on the note. Project + scope is repo-only. Pre-flag notes read as unknown. Run `shadow` and read the + logs before `on`. +- **S24 (PR #52):** host and secret findings cannot fire until the executor or + credential layer reports hosts contacted and secrets injected. Enforcement is + untouched and stays with that layer. +- **S26 (PR #48):** records disagreements only; the sandbox pool still fails a + run whose host dies. +- **S13 (PR #50):** probes are test-only; five operations unprobed. +- **S15/S17 (PR #46):** inert; callers must redact log excerpts. +- **S22 (PR #47):** the coordination record does not keep the revision the + client was planned against. +- **S19 (PR #49):** doctor's store and clock checks report unknown until real + probes exist; the restore comparison has no snapshot producer. +- **S07 (PR #53):** unwired; wiring must not add a durable step to existing runs. +- **Agent hygiene:** parallel agents share one scratch directory and the fixed + grading port 8901. A backup-file collision corrupted one agent's mutation + backup (caught and fixed before push) and spurious grader-test failures + appeared only in agent sandboxes. Later waves told agents to use their own + directories and skip those tests locally. + ## 2026-10-04 — wave 3 wired behind flags (S01/S04, S05/S06, S09, S11, S16, S18, S23, S25) All off by default or observe-only; see the programme's wave-3 table for flags. diff --git a/NEXT_SESSION.md b/NEXT_SESSION.md index 14d9364..741e0ec 100644 --- a/NEXT_SESSION.md +++ b/NEXT_SESSION.md @@ -1,12 +1,35 @@ # Next session -The single forward plan is [docs/SHIP_RELEASE_PROGRAMME_2026-09-21.md](docs/SHIP_RELEASE_PROGRAMME_2026-09-21.md); its **Execution status** section is the only status list. Open audit items and deferred findings live in [AUDIT_OPEN.md](AUDIT_OPEN.md). Update both rather than creating another plan. +The single forward plan is [docs/SHIP_RELEASE_PROGRAMME_2026-09-21.md](docs/SHIP_RELEASE_PROGRAMME_2026-09-21.md); its **Execution status** section (the wave-2, wave-3 and wave-4 tables) is the only status list. Open audit items and deferred findings live in [AUDIT_OPEN.md](AUDIT_OPEN.md). Update both rather than creating another plan. -Suggested next slices, in order, each with its own small specification in the programme before implementation: +## Where things stand (2026-10-04) -1. First-batch item 3 (S03): map existing run/intake/follow-up records to the task/requirements/acceptance contract and rehearse the additive migration on a restored copy. -2. First-batch item 2 (S02): deterministic preflight and canary, then propose the bounded model batch for authorisation. -3. S28: wire `runTiming` (`src/run-timing.ts`) into an authorised surface. -4. S01: credential mediation (needs Sandbox `env` support and a live proof) and the teploy-cli preview network change. +Everything below is **implemented and checked by automated tests only**. No row has a verified outcome: no live Ship, Nucleus, sandbox, forge, preview target or model gateway was reachable, and nothing was spent. Most wave-2 modules are tested but not called by any route or worker; waves 3 and 4 wired many of them behind flags that are **off by default** (or observe-only), so default behaviour is unchanged. + +## Needs a human decision first + +1. **Paid evaluation batch** (S02, propose-only): 66 runs, hard cap $25, details in the programme's S02 slice. It must run from somewhere that can reach a Ship instance. The graders were tightened in PR #34, so the retained results were scored under looser graders and were not re-graded. +2. S28 percentile threshold (n>=20) was an agent's choice. +3. S25 service-account role (shadow treats service accounts as id-only matches). +4. Whether S08 test-integrity findings should reach the PR body and webhook (would change the worker path). +5. The 79 sub-24px compact targets from the UI audit (keep or enlarge). +6. File the upstream reports: Neutron plain-text 404 page (`docs/UI_AUDIT_2026-10-03.md`), and the Teploy CLI feature requests in `AUDIT_OPEN.md`. Nothing has been filed. +7. When each wired flag should be turned on (shadow first, read the logs). + +## Needs live systems or the owner's side + +S01 credential proofs on a real sandbox with a private repo; preview isolation (teploy-cli); end-to-end journeys; real-Nucleus checks; the S03 storage migration rehearsal on a restored copy; real deployment adapters; real placement targets; human observation of dashboard users; re-grading retained runs where their trees exist; shadow-log review before enabling any flag. + +## Safe next work (no live system needed) + +- Turn shadow findings into decisions once logs exist; otherwise wire the remaining pure modules (S15/S17 into delivery and incidents, S07 into the plan-park point without adding a durable step, S19 snapshot producer and restore-check command, S18 worker wiring). +- The untouched or barely started packages: S04, S12, S10 follow-ups (offline and error states, other roles), remaining S07 parts. +- Add `./plan-grounding`, `./deployment-adapter`, `./policy-inheritance`, `./tool-manifest` and `./teploy-adapter` subpath exports only when something imports them. + +## Working rules that paid off + +- One branch and PR per slice; merge only on green CI, pinned to the checked head; resolve conflicts by merging main in, never rewriting history. +- Agents in parallel worktrees: give each its own scratch directory (the shared one caused a backup collision) and tell them to skip `scripts/grader-sensitivity.test.mjs` locally (fixed port 8901); CI runs it. +- Every wiring change: default-off equivalence test, a real-path test with the flag on, a negative control per rule. Before finishing any slice: `pnpm run lint`, `pnpm test`, and for `web/` changes `cd web && pnpm test && pnpm run build`. Install `web` dependencies first (`cd web && pnpm install --frozen-lockfile`) or the deployment-pin script test fails for an environmental reason. diff --git a/docs/SHIP_RELEASE_PROGRAMME_2026-09-21.md b/docs/SHIP_RELEASE_PROGRAMME_2026-09-21.md index 98388cc..3ed61fa 100644 --- a/docs/SHIP_RELEASE_PROGRAMME_2026-09-21.md +++ b/docs/SHIP_RELEASE_PROGRAMME_2026-09-21.md @@ -537,6 +537,24 @@ Wave 3 wires wave-2 modules into the product. **Every change below is off by def | S25 | #43 | `SHIP_POLICY_SHADOW=on` (off) | Policy resolved beside nine existing decision points; disagreements recorded to JSONL; `policy shadow-report` | 20 tests, union-merge control, 3 mutations, default-off byte-identical | No enforcement; agreements not counted (no rate); worker intake budget not hooked; models, tools, retention have no decision point; **service-account role undecided** | | S01, S04 | #44 | `SHIP_SAFE_FETCH` unset = **shadow**, `on`, `off`; `SHIP_SAFE_FETCH_ALLOW` | Pinned fetch is the default for forge API calls in `git.ts` and `forge-state.ts`; shadow warns once per host without changing the request | 17 tests including a real loopback server and a rebinding resolver | Setup route, ladder-steps, worker and deploy entry points; `git clone` and `git fetch` subprocesses unguarded; shadow log must be reviewed on a real deployment before `on`; shadow adds one address lookup per forge call (unmeasured) | +#### Wave 4 (2026-10-04): more wiring, and first parts of the untouched packages + +Same rules as wave 3: every change is off by default or observe-only, with tests comparing default behaviour with and without it. **Implemented and checked only; no verified outcomes, nothing run live, nothing spent.** + +| Package | PR | Flag (default) | Implemented | Checked | Still open (specific) | +| --- | --- | --- | --- | --- | --- | +| S15, S17 | #46 | none (inert modules) | `src/recovery-compat.ts` (migration-aware recovery check: unread or unclassified state holds; a newer contract or irreversible migration is incompatible; a recovery plan never counts as a pass) and `src/incident-evidence.ts` (facts, hypotheses and unknowns kept apart; other-service evidence set aside; fewer than five samples is unknown; confidence capped at medium) | 19 tests, 8 negative controls | Unwired; no migration-state reader; no page timeline; log excerpts clamped not redacted (callers must redact); no live incident | +| S22 | #47 | `SHIP_MISSION_VIEW=on` (off) | Read-only mission view on the coordination page; accepted only when the task record says accepted; waivers shown as waived | 23 tests, 4 negative controls | No launching, persistence or replanning; planned-on revision not recorded on the client child, so a moved upstream with no delivery record is invisible | +| S26 | #48 | `SHIP_PLACEMENT=shadow` (off) | Placement decision recorded beside the sandbox pool's host choice and the host-loss path; `placement shadow-report` | 19 tests through the real pool and durable path, 3 mutations | No enforcement; per-run architecture or browser needs not recorded; no drain concept in the pool; a failed liveness probe looks like a TTL expiry | +| S19 | #49 | none (new command) | `teploy-ship doctor`: install prerequisites with pass, fail or unknown (unknown is never pass), secrets reported by presence only, redacted output; restore-readiness comparison module | 20 tests, 8 negative controls, planted-secret control | Store and clock probes report unknown; no snapshot producer or restore-check command; not in the support bundle; S04 not started | +| S13 | #50 | `SHIP_HARNESS_RECORD=on` (off) | Probes for plan-review, steer and investigate across native, claude-code and opencode (fake binaries); harness identity recorded as a sibling key on run-started (step fingerprint unchanged); "Ran on" block on the run page; credential-lifecycle tests | 15 root and 4 web tests, negative controls | Real vendor binaries; five operations unprobed; probes not run at startup or in CI | +| S21 | #51 | `SHIP_KNOWLEDGE_PROVENANCE=off|shadow|on` (off) | Sidecar provenance store (new table, no migration); notes screened and labelled fresh, stale or unknown in `on`; redaction cascades on delete; summaries recorded as agent claims | 25 tests, 6 negative controls, durable step sequence identical across modes | No live Nucleus; freshness coarse; replay-exact derivation not available; cli.ts and agent.ts paths unwired; project scope is repo-only | +| S24 | #52 | `SHIP_TOOL_MANIFEST=shadow|on`, `SHIP_EVENT_ENVELOPE=on` (both off) | `tool validate` dry-run CLI; manifest conformance findings recorded per tool call (never block); signed event envelope as extra webhook headers; client verification example | 14 tests, 4 negative controls, default-off header and body byte identity | The loop sees only the action kind, so host and secret findings cannot fire until the executor reports them; no tool registry or web view | +| S07 | #53 | `SHIP_PLAN_GROUNDING=on` (off; no caller yet) | `src/plan-grounding.ts`: whether files, symbols, scripts and make targets a plan names exist in the committed tree at a stated revision | 15 tests on real git repos, 7 mutation controls | Not wired; "grounded" means the name exists, not that the plan is right; most of S07 not started | +| S27 | #54 | `SHIP_DEPLOY_ADAPTER=teploy` (off) | `TeployAdapter` and an optional adapter path in `delivery.ts` and `worker.ts`; capabilities declare what it cannot do | 28 tests against a fake CLI as a real process, conformance 22 pass, 5 skipped for declared gaps, 10 mutations | No real teploy run; a manual deploy landing inside Ship's deploy is overwritten and still reads back confirmed (pinned by a test); stricter than legacy (unreadable status holds); S27 acceptance not met | + +Upstream feature requests for the Teploy CLI, **not yet filed**: a deploy dry-run; a deploy generation in `status --json` with a compare-and-set or lock that manual deploys honour; documented `logs` flags and a status that separates "no deployment yet" from "unreachable". + Also merged this wave: #9 plan and inventory, #10 devalue 5.9.3 override. S10 and S02 grader sensitivity landed after this table was first written and are the first two rows after S13/S06.