Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions .env.example
Original file line number Diff line number Diff line change
Expand Up @@ -43,6 +43,8 @@ FOREMAN_VALIDATION_COMMANDS=[]
FOREMAN_FORMAT_COMMAND=
FOREMAN_VALIDATION_TIMEOUT_MS=120000
FOREMAN_VALIDATION_MAX_OUTPUT_BYTES=1048576
# Times a Worker is resumed in its own session with failing checks before its turn ends (0 turns this off).
FOREMAN_WORKER_CHECK_ROUNDS=2
# Every validation command runs inside a bubblewrap (bwrap) sandbox that hides your home
# directory, credentials, Foreman's data dir, the source checkout, /tmp and /run. Validation
# fails if bwrap is missing. Set none only on hosts without bubblewrap (for example macOS);
Expand Down
3 changes: 3 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -65,6 +65,8 @@ Validation commands run in a disposable copy of the Worker workspace. Every conf

**Checks that also fail on the base commit.** When validation fails on a Worker snapshot that has changes, Foreman runs the configured commands up to the last failed check (so installs and builds a check depends on run too) once against the unchanged pinned base commit, in the same sandbox, and caches the result on the run (pinned base + command digest). Each failed check is marked `failsOnBase` and shown with a "Fails on base" badge. If every failed check also fails there (for example `pnpm run smoke:install` with no network), a Worker retry cannot fix it, so Foreman does not spend another Worker attempt: the run stops with "Checks also fail on the base commit, so a Worker retry cannot fix them: ...". Fix the check or its network setting, then use **Retry validation** (which also discards the cached baseline). If some failures are Worker-caused, the normal automatic follow-up runs, but its correction note lists only those and names the base-failing checks as not the Worker's to fix. This never relaxes promotion: every check must still pass. If the baseline cannot run, Foreman falls back to the ordinary follow-up.

**Worker check gate.** Workers are interchangeable, so none of them is trusted to run the repository's checks: Antigravity and Claude Code have file tools only, and Codex's sandbox has no network and no installed dependencies. Instead, when a Worker's turn completes the bridge keeps the task in progress and pauses it, and Foreman runs the configured checks on the paused workspace (the same snapshot verification, format step and sandboxed validation as the final pass). If a check the Worker could have caused fails, Foreman sends the command and a head-and-tail output excerpt back, and the bridge resumes the same CLI session (`--resume` for Claude Code, `exec resume` for Codex, `--conversation` for Antigravity) so the Worker fixes it with its full context. A change outside the allowed scope counts as a failure; checks that fail on the base commit and any infrastructure problem do not. This repeats up to `FOREMAN_WORKER_CHECK_ROUNDS` times (default 2) within one Worker attempt, so it spends no Worker attempt and no Orchestrator turn. The gate's verdicts are feedback only: Foreman's validation after the turn is still the result the Reviewer and you see. It is off when automatic validation correction is off for the run. Each round is recorded as a `worker.check_round` event and in the bridge response's `metadata.worker_gate`.

**Format step (optional).** Antigravity and Claude Code Workers can only edit files, so they cannot run the repository's formatter, and a `prettier --check` validation command would fail on their output. When a `formatCommand` is configured (`{"name","command","args","cwd?","network?"}`; the open-repository dialog suggests `<runner> run format`, `format:write` or `prettier:write` from `package.json` with an on/off toggle, `FOREMAN_FORMAT_COMMAND` sets it for the single-repository configuration), Foreman runs it after it has verified the Worker snapshot and before validation. It materializes the verified snapshot in the same bubblewrap sandbox as validation, runs the dependency installs from the validation list (`pnpm`, `npm` or `yarn` install, `ci` or `add`), then the formatter (offline unless it sets `network: true`). Only the files the Worker added, modified or renamed are read back, keeping the Worker's file mode; anything else the formatter touched or created is ignored. The new snapshot is re-verified against the pinned base and allowed scope, and validation, the Reviewer, approval and promotion all use the formatted bytes. The run's evidence records `formatting` (status `applied`, `unchanged` or `failed`, the formatted paths and the formatter's bounded output) and the UI shows "Foreman formatted N files". A formatter failure never fails the run: the Worker's snapshot is kept and validation reports the real problem. The step applies to live Worker snapshots only, not to recorded replays.

The suggested allowed scope covers the top-level tracked paths except CI configuration (`.github/`, `.gitlab-ci.yml`, `.circleci/`, `.buildkite/`, `azure-pipelines.yml`, `Jenkinsfile`, `.travis.yml`), which runs with repository secrets once pushed. Add those paths by hand if a task really has to change them.
Expand Down Expand Up @@ -155,6 +157,7 @@ The standard local flow requires no environment variables. The following setting
| `FOREMAN_FORMAT_COMMAND` | — | Optional JSON object `{"name","command","args","cwd?","network?"}`: the formatter Foreman runs on the Worker's changed files before validation. Offline unless `network` is `true`. See the format step above. |
| `FOREMAN_VALIDATION_TIMEOUT_MS` | `120000` | Per-command time limit (max 600,000 ms). |
| `FOREMAN_VALIDATION_MAX_OUTPUT_BYTES` | `1048576` | Per-command output capture bound (max 16 MiB). |
| `FOREMAN_WORKER_CHECK_ROUNDS` | `2` | Worker check gate: how many times a Worker is resumed in its own session with failing checks before its turn ends (max 3). `0` turns the gate off. See the check gate above. |
| `FOREMAN_VALIDATION_SANDBOX` | `bwrap` | `bwrap` runs every validation command in a bubblewrap sandbox and fails if it is unavailable. `none` runs them directly on the host with your credentials, and is unsafe. See [Validation sandbox](#validation-sandbox). |
| `FOREMAN_VALIDATION_SANDBOX_RO_PATHS` | — | Comma-separated absolute paths mounted read-only in the sandbox, for toolchains under a hidden directory such as `$HOME/.volta`. Paths that do not exist are ignored. |

Expand Down
1 change: 1 addition & 0 deletions config.schema.json
Original file line number Diff line number Diff line change
Expand Up @@ -23,6 +23,7 @@
"FOREMAN_FORMAT_COMMAND": { "type": "string", "description": "Optional JSON object {name, command, args, cwd?, network?} for a formatter Foreman runs in the sandboxed validation workspace (after the install commands) on the Worker's added or modified files, before validation. The formatted bytes become the evidence that validation, the Reviewer and promotion use. Runs without network access unless network is true" },
"FOREMAN_VALIDATION_TIMEOUT_MS": { "type": "integer", "minimum": 1, "maximum": 600000, "default": 120000 },
"FOREMAN_VALIDATION_MAX_OUTPUT_BYTES": { "type": "integer", "minimum": 1, "maximum": 16777216, "default": 1048576 },
"FOREMAN_WORKER_CHECK_ROUNDS": { "type": "integer", "minimum": 0, "maximum": 3, "default": 2, "description": "Worker check gate: after a Worker turn, Foreman runs the configured checks on the paused workspace and resumes the same Worker session with any failures, up to this many times. 0 turns the gate off" },
"FOREMAN_VALIDATION_SANDBOX": { "type": "string", "enum": ["bwrap", "none"], "default": "bwrap", "description": "bwrap runs every validation command inside a bubblewrap sandbox and fails closed when it is unavailable; none runs them unsandboxed on the host (unsafe opt-out)" },
"FOREMAN_VALIDATION_SANDBOX_RO_PATHS": { "type": "string", "description": "Comma-separated absolute paths (for example a toolchain under your home directory) mounted read-only into the validation sandbox" }
}
Expand Down
6 changes: 6 additions & 0 deletions investigations/local-cli-uhp/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -234,6 +234,12 @@ idempotency key, and a persisted workspace ID for replay.
`LOCAL_CLI_UHP_SMOKE_KEY_FILE`
and `LOCAL_CLI_UHP_EVIDENCE_FILE` select the key and evidence paths.

## Worker check gate

A Worker request may carry `metadata.foreman_worker_gate: {"max_rounds": 1-3}` (Reviewer and Planner/Orchestrator requests are refused with `worker_gate_role_unsupported`). Discovery advertises it as `extensions.foreman_worker_gate_v1`.

After each completed Worker turn the task stays `in_progress` and the bridge records a `response.activity` event with `kind: "worker_gate"` and `gate_round: N`. While paused, no CLI is running, so the workspace snapshot endpoint serves a complete snapshot for Foreman's checks (overlay stays refused). Foreman answers with `POST /extensions/foreman-workspace/v1/responses/{id}/worker-gate` and `{"round": N, "status": "passed" | "failed" | "skipped", "feedback"?, "failed_checks"?}`; a verdict for any other round is refused with 409. On `failed`, the bridge resumes the same native session with the feedback: `--resume` for Claude Code (the transcript directories are bound per workspace, like a role session), `exec resume` without `--ephemeral` for Codex (its per-workspace `CODEX_HOME` is kept), and `--conversation` for Antigravity (its per-workspace state directory is kept). Any other verdict, no verdict within `LOCAL_CLI_UHP_WORKER_GATE_TIMEOUT_MS` (default 15 minutes), a failed turn or a cancellation ends the task as the last turn left it. `metadata.worker_gate` records each round; usage is summed across Claude Code and Codex turns and taken from the last Antigravity turn, whose usage is cumulative per conversation. `cli_invocation` stays the first turn's invocation. The kept session state is deleted when the task ends.

## Recorded live proof

The first Claude workspace attempt failed closed at `boundary_probe` before CLI
Expand Down
5 changes: 3 additions & 2 deletions investigations/local-cli-uhp/cli-args.mjs
Original file line number Diff line number Diff line change
Expand Up @@ -4,8 +4,9 @@ export function claudeCodeCliArgs(model, { reviewer = false, sessionId, persiste
return ['-p', '--output-format', 'stream-json', '--verbose', '--model', model, '--max-turns', String(Math.min(maxStep, 10)), '--restricted', '--strict-mcp-config', '--permission-mode', persistentContext ? 'plan' : 'acceptEdits', '--tools', persistentContext ? 'Read,Grep,Glob' : 'Read,Edit,Write', ...(sessionId ? ['--resume', sessionId] : [])];
}

export function codexCliArgs(model, { reviewer = false, sessionId, persistentContext = false } = {}) {
/** `keepSession` keeps a Worker's session on disk (no --ephemeral) so a Worker check round can resume it. */
export function codexCliArgs(model, { reviewer = false, sessionId, persistentContext = false, keepSession = false } = {}) {
// --sandbox is a top-level Codex option. `codex exec resume` rejects it when
// placed after `resume`, before any session can be reported.
return ['--ask-for-approval', 'never', '--sandbox', reviewer || persistentContext ? 'read-only' : 'workspace-write', 'exec', ...(sessionId ? ['resume', sessionId] : []), '--json', ...(!persistentContext ? ['--ephemeral'] : []), '--ignore-user-config', ...(reviewer ? ['--ignore-rules'] : []), '--skip-git-repo-check', '--model', model, '-'];
return ['--ask-for-approval', 'never', '--sandbox', reviewer || persistentContext ? 'read-only' : 'workspace-write', 'exec', ...(sessionId ? ['resume', sessionId] : []), '--json', ...(!persistentContext && !keepSession ? ['--ephemeral'] : []), '--ignore-user-config', ...(reviewer ? ['--ignore-rules'] : []), '--skip-git-repo-check', '--model', model, '-'];
}
Loading
Loading