Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
32 changes: 32 additions & 0 deletions AUDIT_OPEN.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,37 @@
# Open audit items

## 2026-10-04 — wave 4 (S07, S13, S15/S17, S19, S21, S22, S24, S26, S27)

All off by default or observe-only; flags are in the programme's wave-4 table.

- **S27 Teploy adapter (PR #54):** flag off. When on, it is stricter than the
legacy path (unreadable status, including a brand-new destination, holds).
A manual `teploy deploy` landing inside Ship's deploy is overwritten and still
reads back confirmed; the lease serialises Ship's own releases on one host
only. Upstream feature requests (not filed): deploy dry-run; deploy generation
and compare-and-set in `status --json`; documented `logs` flags and a
no-deployment-versus-unreachable signal.
- **S21 (PR #51):** provenance is a sidecar, not fields on the note. Project
scope is repo-only. Pre-flag notes read as unknown. Run `shadow` and read the
logs before `on`.
- **S24 (PR #52):** host and secret findings cannot fire until the executor or
credential layer reports hosts contacted and secrets injected. Enforcement is
untouched and stays with that layer.
- **S26 (PR #48):** records disagreements only; the sandbox pool still fails a
run whose host dies.
- **S13 (PR #50):** probes are test-only; five operations unprobed.
- **S15/S17 (PR #46):** inert; callers must redact log excerpts.
- **S22 (PR #47):** the coordination record does not keep the revision the
client was planned against.
- **S19 (PR #49):** doctor's store and clock checks report unknown until real
probes exist; the restore comparison has no snapshot producer.
- **S07 (PR #53):** unwired; wiring must not add a durable step to existing runs.
- **Agent hygiene:** parallel agents share one scratch directory and the fixed
grading port 8901. A backup-file collision corrupted one agent's mutation
backup (caught and fixed before push) and spurious grader-test failures
appeared only in agent sandboxes. Later waves told agents to use their own
directories and skip those tests locally.

## 2026-10-04 — wave 3 wired behind flags (S01/S04, S05/S06, S09, S11, S16, S18, S23, S25)

All off by default or observe-only; see the programme's wave-3 table for flags.
Expand Down
35 changes: 29 additions & 6 deletions NEXT_SESSION.md
Original file line number Diff line number Diff line change
@@ -1,12 +1,35 @@
# Next session

The single forward plan is [docs/SHIP_RELEASE_PROGRAMME_2026-09-21.md](docs/SHIP_RELEASE_PROGRAMME_2026-09-21.md); its **Execution status** section is the only status list. Open audit items and deferred findings live in [AUDIT_OPEN.md](AUDIT_OPEN.md). Update both rather than creating another plan.
The single forward plan is [docs/SHIP_RELEASE_PROGRAMME_2026-09-21.md](docs/SHIP_RELEASE_PROGRAMME_2026-09-21.md); its **Execution status** section (the wave-2, wave-3 and wave-4 tables) is the only status list. Open audit items and deferred findings live in [AUDIT_OPEN.md](AUDIT_OPEN.md). Update both rather than creating another plan.

Suggested next slices, in order, each with its own small specification in the programme before implementation:
## Where things stand (2026-10-04)

1. First-batch item 3 (S03): map existing run/intake/follow-up records to the task/requirements/acceptance contract and rehearse the additive migration on a restored copy.
2. First-batch item 2 (S02): deterministic preflight and canary, then propose the bounded model batch for authorisation.
3. S28: wire `runTiming` (`src/run-timing.ts`) into an authorised surface.
4. S01: credential mediation (needs Sandbox `env` support and a live proof) and the teploy-cli preview network change.
Everything below is **implemented and checked by automated tests only**. No row has a verified outcome: no live Ship, Nucleus, sandbox, forge, preview target or model gateway was reachable, and nothing was spent. Most wave-2 modules are tested but not called by any route or worker; waves 3 and 4 wired many of them behind flags that are **off by default** (or observe-only), so default behaviour is unchanged.

## Needs a human decision first

1. **Paid evaluation batch** (S02, propose-only): 66 runs, hard cap $25, details in the programme's S02 slice. It must run from somewhere that can reach a Ship instance. The graders were tightened in PR #34, so the retained results were scored under looser graders and were not re-graded.
2. S28 percentile threshold (n>=20) was an agent's choice.
3. S25 service-account role (shadow treats service accounts as id-only matches).
4. Whether S08 test-integrity findings should reach the PR body and webhook (would change the worker path).
5. The 79 sub-24px compact targets from the UI audit (keep or enlarge).
6. File the upstream reports: Neutron plain-text 404 page (`docs/UI_AUDIT_2026-10-03.md`), and the Teploy CLI feature requests in `AUDIT_OPEN.md`. Nothing has been filed.
7. When each wired flag should be turned on (shadow first, read the logs).

## Needs live systems or the owner's side

S01 credential proofs on a real sandbox with a private repo; preview isolation (teploy-cli); end-to-end journeys; real-Nucleus checks; the S03 storage migration rehearsal on a restored copy; real deployment adapters; real placement targets; human observation of dashboard users; re-grading retained runs where their trees exist; shadow-log review before enabling any flag.

## Safe next work (no live system needed)

- Turn shadow findings into decisions once logs exist; otherwise wire the remaining pure modules (S15/S17 into delivery and incidents, S07 into the plan-park point without adding a durable step, S19 snapshot producer and restore-check command, S18 worker wiring).
- The untouched or barely started packages: S04, S12, S10 follow-ups (offline and error states, other roles), remaining S07 parts.
- Add `./plan-grounding`, `./deployment-adapter`, `./policy-inheritance`, `./tool-manifest` and `./teploy-adapter` subpath exports only when something imports them.

## Working rules that paid off

- One branch and PR per slice; merge only on green CI, pinned to the checked head; resolve conflicts by merging main in, never rewriting history.
- Agents in parallel worktrees: give each its own scratch directory (the shared one caused a backup collision) and tell them to skip `scripts/grader-sensitivity.test.mjs` locally (fixed port 8901); CI runs it.
- Every wiring change: default-off equivalence test, a real-path test with the flag on, a negative control per rule.

Before finishing any slice: `pnpm run lint`, `pnpm test`, and for `web/` changes `cd web && pnpm test && pnpm run build`. Install `web` dependencies first (`cd web && pnpm install --frozen-lockfile`) or the deployment-pin script test fails for an environmental reason.
18 changes: 18 additions & 0 deletions docs/SHIP_RELEASE_PROGRAMME_2026-09-21.md
Original file line number Diff line number Diff line change
Expand Up @@ -537,6 +537,24 @@ Wave 3 wires wave-2 modules into the product. **Every change below is off by def
| S25 | #43 | `SHIP_POLICY_SHADOW=on` (off) | Policy resolved beside nine existing decision points; disagreements recorded to JSONL; `policy shadow-report` | 20 tests, union-merge control, 3 mutations, default-off byte-identical | No enforcement; agreements not counted (no rate); worker intake budget not hooked; models, tools, retention have no decision point; **service-account role undecided** |
| S01, S04 | #44 | `SHIP_SAFE_FETCH` unset = **shadow**, `on`, `off`; `SHIP_SAFE_FETCH_ALLOW` | Pinned fetch is the default for forge API calls in `git.ts` and `forge-state.ts`; shadow warns once per host without changing the request | 17 tests including a real loopback server and a rebinding resolver | Setup route, ladder-steps, worker and deploy entry points; `git clone` and `git fetch` subprocesses unguarded; shadow log must be reviewed on a real deployment before `on`; shadow adds one address lookup per forge call (unmeasured) |

#### Wave 4 (2026-10-04): more wiring, and first parts of the untouched packages

Same rules as wave 3: every change is off by default or observe-only, with tests comparing default behaviour with and without it. **Implemented and checked only; no verified outcomes, nothing run live, nothing spent.**

| Package | PR | Flag (default) | Implemented | Checked | Still open (specific) |
| --- | --- | --- | --- | --- | --- |
| S15, S17 | #46 | none (inert modules) | `src/recovery-compat.ts` (migration-aware recovery check: unread or unclassified state holds; a newer contract or irreversible migration is incompatible; a recovery plan never counts as a pass) and `src/incident-evidence.ts` (facts, hypotheses and unknowns kept apart; other-service evidence set aside; fewer than five samples is unknown; confidence capped at medium) | 19 tests, 8 negative controls | Unwired; no migration-state reader; no page timeline; log excerpts clamped not redacted (callers must redact); no live incident |
| S22 | #47 | `SHIP_MISSION_VIEW=on` (off) | Read-only mission view on the coordination page; accepted only when the task record says accepted; waivers shown as waived | 23 tests, 4 negative controls | No launching, persistence or replanning; planned-on revision not recorded on the client child, so a moved upstream with no delivery record is invisible |
| S26 | #48 | `SHIP_PLACEMENT=shadow` (off) | Placement decision recorded beside the sandbox pool's host choice and the host-loss path; `placement shadow-report` | 19 tests through the real pool and durable path, 3 mutations | No enforcement; per-run architecture or browser needs not recorded; no drain concept in the pool; a failed liveness probe looks like a TTL expiry |
| S19 | #49 | none (new command) | `teploy-ship doctor`: install prerequisites with pass, fail or unknown (unknown is never pass), secrets reported by presence only, redacted output; restore-readiness comparison module | 20 tests, 8 negative controls, planted-secret control | Store and clock probes report unknown; no snapshot producer or restore-check command; not in the support bundle; S04 not started |
| S13 | #50 | `SHIP_HARNESS_RECORD=on` (off) | Probes for plan-review, steer and investigate across native, claude-code and opencode (fake binaries); harness identity recorded as a sibling key on run-started (step fingerprint unchanged); "Ran on" block on the run page; credential-lifecycle tests | 15 root and 4 web tests, negative controls | Real vendor binaries; five operations unprobed; probes not run at startup or in CI |
| S21 | #51 | `SHIP_KNOWLEDGE_PROVENANCE=off|shadow|on` (off) | Sidecar provenance store (new table, no migration); notes screened and labelled fresh, stale or unknown in `on`; redaction cascades on delete; summaries recorded as agent claims | 25 tests, 6 negative controls, durable step sequence identical across modes | No live Nucleus; freshness coarse; replay-exact derivation not available; cli.ts and agent.ts paths unwired; project scope is repo-only |
| S24 | #52 | `SHIP_TOOL_MANIFEST=shadow|on`, `SHIP_EVENT_ENVELOPE=on` (both off) | `tool validate` dry-run CLI; manifest conformance findings recorded per tool call (never block); signed event envelope as extra webhook headers; client verification example | 14 tests, 4 negative controls, default-off header and body byte identity | The loop sees only the action kind, so host and secret findings cannot fire until the executor reports them; no tool registry or web view |
| S07 | #53 | `SHIP_PLAN_GROUNDING=on` (off; no caller yet) | `src/plan-grounding.ts`: whether files, symbols, scripts and make targets a plan names exist in the committed tree at a stated revision | 15 tests on real git repos, 7 mutation controls | Not wired; "grounded" means the name exists, not that the plan is right; most of S07 not started |
| S27 | #54 | `SHIP_DEPLOY_ADAPTER=teploy` (off) | `TeployAdapter` and an optional adapter path in `delivery.ts` and `worker.ts`; capabilities declare what it cannot do | 28 tests against a fake CLI as a real process, conformance 22 pass, 5 skipped for declared gaps, 10 mutations | No real teploy run; a manual deploy landing inside Ship's deploy is overwritten and still reads back confirmed (pinned by a test); stricter than legacy (unreadable status holds); S27 acceptance not met |

Upstream feature requests for the Teploy CLI, **not yet filed**: a deploy dry-run; a deploy generation in `status --json` with a compare-and-set or lock that manual deploys honour; documented `logs` flags and a status that separates "no deployment yet" from "unreachable".


Also merged this wave: #9 plan and inventory, #10 devalue 5.9.3 override. S10 and S02 grader sensitivity landed after this table was first written and are the first two rows after S13/S06.

Expand Down
Loading