Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 7 additions & 0 deletions .env.example
Original file line number Diff line number Diff line change
Expand Up @@ -34,6 +34,13 @@ FOREMAN_WORKSPACE_ALLOWED_SCOPE=
# Without the field, a pnpm/npm/yarn install (install, i, ci, add, or bare yarn) gets the
# network and every other command does not.
FOREMAN_VALIDATION_COMMANDS=[]
# Optional JSON object, same shape as one validation command. Foreman runs it in the sandboxed
# validation workspace (after the install commands above) on the Worker's changed files, before
# validation, and validates, reviews and promotes the formatted bytes. Use it for repositories whose
# checks include a formatter (for example prettier --check) that file-only Workers cannot run.
# Only files the Worker added or modified are kept. It runs without network unless "network": true.
# Example: {"name":"Format","command":"pnpm","args":["run","format"]}
FOREMAN_FORMAT_COMMAND=
FOREMAN_VALIDATION_TIMEOUT_MS=120000
FOREMAN_VALIDATION_MAX_OUTPUT_BYTES=1048576
# Every validation command runs inside a bubblewrap (bwrap) sandbox that hides your home
Expand Down
7 changes: 6 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -61,6 +61,10 @@ When you open a repository, Foreman inspects `.github/workflows` and suggests ma

Validation commands run in a disposable copy of the Worker workspace. Every configured command must pass. Commands absent from the list are not run or inferred.

**Checks that also fail on the base commit.** When validation fails on a Worker snapshot that has changes, Foreman runs the configured commands up to the last failed check (so installs and builds a check depends on run too) once against the unchanged pinned base commit, in the same sandbox, and caches the result on the run (pinned base + command digest). Each failed check is marked `failsOnBase` and shown with a "Fails on base" badge. If every failed check also fails there (for example `pnpm run smoke:install` with no network), a Worker retry cannot fix it, so Foreman does not spend another Worker attempt: the run stops with "Checks also fail on the base commit, so a Worker retry cannot fix them: ...". Fix the check or its network setting, then use **Retry validation** (which also discards the cached baseline). If some failures are Worker-caused, the normal automatic follow-up runs, but its correction note lists only those and names the base-failing checks as not the Worker's to fix. This never relaxes promotion: every check must still pass. If the baseline cannot run, Foreman falls back to the ordinary follow-up.

**Format step (optional).** Antigravity and Claude Code Workers can only edit files, so they cannot run the repository's formatter, and a `prettier --check` validation command would fail on their output. When a `formatCommand` is configured (`{"name","command","args","cwd?","network?"}`; the open-repository dialog suggests `<runner> run format`, `format:write` or `prettier:write` from `package.json` with an on/off toggle, `FOREMAN_FORMAT_COMMAND` sets it for the single-repository configuration), Foreman runs it after it has verified the Worker snapshot and before validation. It materializes the verified snapshot in the same bubblewrap sandbox as validation, runs the dependency installs from the validation list (`pnpm`, `npm` or `yarn` install, `ci` or `add`), then the formatter (offline unless it sets `network: true`). Only the files the Worker added, modified or renamed are read back, keeping the Worker's file mode; anything else the formatter touched or created is ignored. The new snapshot is re-verified against the pinned base and allowed scope, and validation, the Reviewer, approval and promotion all use the formatted bytes. The run's evidence records `formatting` (status `applied`, `unchanged` or `failed`, the formatted paths and the formatter's bounded output) and the UI shows "Foreman formatted N files". A formatter failure never fails the run: the Worker's snapshot is kept and validation reports the real problem. The step applies to live Worker snapshots only, not to recorded replays.

The suggested allowed scope covers the top-level tracked paths except CI configuration (`.github/`, `.gitlab-ci.yml`, `.circleci/`, `.buildkite/`, `azure-pipelines.yml`, `Jenkinsfile`, `.travis.yml`), which runs with repository secrets once pushed. Add those paths by hand if a task really has to change them.

The checks pipeline in the review panel shows local Foreman results alongside remote GitHub CI status. CI failure excerpts are fetched from `GET /api/runs/:id/github/ci-failures`.
Expand All @@ -75,7 +79,7 @@ The Worker's files are untrusted, and validation runs them (`pnpm install` lifec
- **Package-manager cache:** `<data dir>/validation-cache` (owner-only) is mounted read-write at `/tmp/foreman-cache`, with `XDG_CACHE_HOME`, `XDG_DATA_HOME`, `npm_config_cache`, `npm_config_store_dir` and `pnpm_config_store_dir` (pnpm 10 and 11), `YARN_CACHE_FOLDER` and `COREPACK_HOME` pointing into it, so installs do not download everything each time. It is shared by every validation and project. pnpm and npm verify package integrity against the lockfile, but a command that ran hostile code could still leave bad data behind, for example a swapped pnpm binary under `XDG_DATA_HOME`. Delete the directory if you suspect that.
- **Network policy: none, except commands that explicitly need it.** Every validation command runs in an empty network namespace (`bwrap --unshare-net`). It has no interface except a working loopback, so tests that bind and connect to `127.0.0.1` inside the sandbox still work, but it cannot reach Foreman's own API, the per-project bridge or any other service on the host's loopback, the host's abstract-namespace Unix sockets, the cloud instance metadata endpoint, the local network or the internet, and DNS lookups fail. A hostile test therefore cannot drive Foreman's API for another run or fetch instance credentials. If bubblewrap cannot create the network namespace on your host (some containers refuse), the sandbox probe fails and validation reports the sandbox as unavailable; it never falls back to the host network.
- **The `network` flag.** Every validation command takes an optional boolean `network`: in `FOREMAN_VALIDATION_COMMANDS` (`{"name","command","args","cwd?","network?"}`), in a saved workspace setup, in the task-start command overrides, and as the per-command **Network** checkbox in the open-repository dialog and the task's advanced options. Only `true` keeps the host network for that command; only the JSON values `true` and `false` are accepted, anything else is rejected. Each check records what it got as `network` (with `sandbox: none` every check is recorded as `network: true`, since nothing isolates it), and the checks pipeline shows a **Network** badge on checks that ran with network access.
- **Which commands get the network by default.** The open-repository dialog starts each command from the repository inspector's suggestion: the dependency install (`pnpm install --frozen-lockfile`, `npm ci`) has `network: true`, and so do `cargo test` and `go test ./...` because they download dependencies on first run. Every other suggestion (`pnpm run test`, `pnpm run typecheck`, `python -m pytest`, ...) is offline; the install step puts `node_modules` in the shared validation workspace, so the checks after it need no network. A command that fetches tooling implicitly (for example pnpm downloading the version pinned in `packageManager`) fails offline and needs the checkbox.
- **Which commands get the network by default.** The open-repository dialog starts each command from the repository inspector's suggestion: the dependency install (`pnpm install --frozen-lockfile`, `npm ci`) has `network: true`, and so do `cargo test` and `go test ./...` because they download dependencies on first run. A package script also gets `network: true` when it is a smoke script (`smoke`, `smoke:install`, `smoke:load`, ...), is named for an install (`...:install`), or its body runs a package install, `pnpm dlx` or `npx`, since it would fail offline whatever the Worker changed. Every other suggestion (`pnpm run test`, `pnpm run typecheck`, `python -m pytest`, ...) is offline; the install step puts `node_modules` in the shared validation workspace, so the checks after it need no network. A command that fetches tooling implicitly (for example pnpm downloading the version pinned in `packageManager`) fails offline and needs the checkbox.
- **Older setups and configs.** A command with no `network` field, in a saved workspace setup, in `FOREMAN_VALIDATION_COMMANDS` or in a task-start override, is treated as `network: true` if it is a recognised package-manager install and `network: false` otherwise. Recognised means the executable is exactly `pnpm`, `npm` or `yarn` with first argument `install`, `i`, `ci` or `add`, or bare `yarn` with no arguments. An explicit `network` always wins, so `{"command":"pnpm","args":["install","--offline"],"network":false}` stays offline. Runs stored before the flag existed are not rewritten (their approval is bound to a digest of the stored command list); a stored command without the field gets the same default when it runs, and the check records the effective value.
- **Residual risk, network-enabled commands.** A command with `network: true` shares the host network namespace, including the host's loopback services and the metadata endpoint. Package lifecycle scripts (`preinstall`, `postinstall`, `prepare`) run during a network-enabled install, so a hostile dependency or Worker-edited manifest gets the network for the length of that install. Grant `network` only to the install step, and consider `--ignore-scripts` if your dependencies allow it. The offline default protects every other command, including the tests.
- Run Foreman as an unprivileged user. The sandbox hides paths but cannot stop a root process from reading root-only files that remain visible, such as `/etc/shadow`.
Expand Down Expand Up @@ -146,6 +150,7 @@ The standard local flow requires no environment variables. The following setting
| `FOREMAN_WORKSPACE_BRIDGE_TOKEN` | — | Optional bearer token for that bridge, for a bridge started with `LOCAL_CLI_UHP_TOKEN` (use the same value). Requires `FOREMAN_WORKSPACE_BRIDGE_URL`; printable ASCII without spaces. It is only ever sent to that loopback URL and is never logged or returned by the API. |
| `FOREMAN_WORKSPACE_ALLOWED_SCOPE` | — | Comma-separated exact paths or directory prefixes ending in `/` allowed in Worker results. |
| `FOREMAN_VALIDATION_COMMANDS` | — | JSON array of `{"name","command","args","cwd?","network?"}` entries run in the disposable validation workspace. `network` is `true` or `false`; validation runs offline unless it is `true`, and a package-manager install without the field defaults to `true`. See [Validation sandbox](#validation-sandbox). |
| `FOREMAN_FORMAT_COMMAND` | — | Optional JSON object `{"name","command","args","cwd?","network?"}`: the formatter Foreman runs on the Worker's changed files before validation. Offline unless `network` is `true`. See the format step above. |
| `FOREMAN_VALIDATION_TIMEOUT_MS` | `120000` | Per-command time limit (max 600,000 ms). |
| `FOREMAN_VALIDATION_MAX_OUTPUT_BYTES` | `1048576` | Per-command output capture bound (max 16 MiB). |
| `FOREMAN_VALIDATION_SANDBOX` | `bwrap` | `bwrap` runs every validation command in a bubblewrap sandbox and fails if it is unavailable. `none` runs them directly on the host with your credentials, and is unsafe. See [Validation sandbox](#validation-sandbox). |
Expand Down
1 change: 1 addition & 0 deletions config.schema.json
Original file line number Diff line number Diff line change
Expand Up @@ -20,6 +20,7 @@
"FOREMAN_WORKSPACE_BRIDGE_TOKEN": { "type": "string", "writeOnly": true, "minLength": 1, "maxLength": 512, "pattern": "^[\\x21-\\x7e]+$", "description": "Optional bearer token for the external workspace bridge (its LOCAL_CLI_UHP_TOKEN); printable ASCII without spaces; requires FOREMAN_WORKSPACE_BRIDGE_URL" },
"FOREMAN_WORKSPACE_ALLOWED_SCOPE": { "type": "string", "description": "Comma-separated exact paths or directory prefixes ending in / allowed for Worker changes" },
"FOREMAN_VALIDATION_COMMANDS": { "type": "string", "description": "JSON array of validation command objects {name, command, args, cwd?, network?} with explicit command and args. network is true or false (no other value is accepted): validation commands run without network access unless it is true. When omitted, a pnpm, npm or yarn install (first argument install, i, ci or add, or bare yarn) defaults to true and every other command to false" },
"FOREMAN_FORMAT_COMMAND": { "type": "string", "description": "Optional JSON object {name, command, args, cwd?, network?} for a formatter Foreman runs in the sandboxed validation workspace (after the install commands) on the Worker's added or modified files, before validation. The formatted bytes become the evidence that validation, the Reviewer and promotion use. Runs without network access unless network is true" },
"FOREMAN_VALIDATION_TIMEOUT_MS": { "type": "integer", "minimum": 1, "maximum": 600000, "default": 120000 },
"FOREMAN_VALIDATION_MAX_OUTPUT_BYTES": { "type": "integer", "minimum": 1, "maximum": 16777216, "default": 1048576 },
"FOREMAN_VALIDATION_SANDBOX": { "type": "string", "enum": ["bwrap", "none"], "default": "bwrap", "description": "bwrap runs every validation command inside a bubblewrap sandbox and fails closed when it is unavailable; none runs them unsandboxed on the host (unsafe opt-out)" },
Expand Down
51 changes: 51 additions & 0 deletions src/baseline-validation.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,51 @@
import { createHash } from 'node:crypto';
import { snapshotGitCommit } from './git-workspace.js';
import { type ValidationSandboxConfig } from './validation-sandbox.js';
import { validateWorkerOutput, type ValidationCommand, type VerifiedWorkerWorkspace } from './verified-workspace.js';
import type { BaselineValidation, ValidationCheck } from './domain.js';

/** Identity of a configured command list. Any change to a command's name, argv, cwd or network setting invalidates a cached baseline. */
export const baselineCommandDigest=(commands:readonly ValidationCommand[]):string=>createHash('sha256').update(JSON.stringify(commands.map(c=>[c.name,c.command,c.args,c.cwd??null,c.network??null]))).digest('hex');

/** The commands a baseline run needs, in configured order: every command up to the last named check, since a check can depend on any earlier step (an install, a build), not only on setup commands. */
export const baselineCommandsFor=(commands:readonly ValidationCommand[],names:ReadonlySet<string>):ValidationCommand[]=>{let last=-1;commands.forEach((c,i)=>{if(names.has(c.name))last=i;});return commands.slice(0,last+1);};
const passedCheck=(c:{exitCode:number|null;timedOut:boolean;outputTruncated:boolean})=>c.exitCode===0&&!c.timedOut&&!c.outputTruncated;

/** The unchanged pinned base as verified evidence with zero changes, so `validateWorkerOutput` materializes and sandboxes it exactly like a Worker snapshot. */
async function unchangedBaseEvidence(repoPath:string,pinnedBaseCommit:string,allowedScope:readonly string[]):Promise<VerifiedWorkerWorkspace>{
const snapshot=await snapshotGitCommit(repoPath,pinnedBaseCommit);
return {provenance:'recorded_replay',pinnedBaseCommit:snapshot.commit,completeSnapshot:{reportedComplete:true,reportedErrors:0,entryCount:snapshot.entries.length},scopeVerified:true,allowedScope:[...allowedScope],entries:snapshot.entries,changes:[],reviewDiff:''};
}

/**
* Run the setup commands plus the named failed checks once against the unchanged pinned base. The result is cached on the run by pinned base + command digest;
* a check already in the cache is never run on the base again, so with an unchanged set of failing checks the baseline runs once per run.
* Only a later failure of a check the cache has not seen extends it (its prerequisites rerun to prepare that workspace).
*/
export async function ensureBaseline(input:{repoPath:string;pinnedBaseCommit:string;allowedScope:readonly string[];commands:readonly ValidationCommand[];failedNames:readonly string[];cached?:BaselineValidation;timeoutMs?:number;maxOutputBytes?:number;sandbox?:ValidationSandboxConfig}):Promise<{baseline:BaselineValidation;ran:string[]}>{
const commandDigest=baselineCommandDigest(input.commands),cached=input.cached&&input.cached.pinnedBaseCommit===input.pinnedBaseCommit&&input.cached.commandDigest===commandDigest?input.cached:undefined;
const known=new Set((cached?.checks??[]).map(c=>c.name)),missing=new Set(input.failedNames.filter(n=>!known.has(n)));
if(cached&&!missing.size)return {baseline:cached,ran:[]};
const commands=baselineCommandsFor(input.commands,missing);
const observed=await validateWorkerOutput({repoPath:input.repoPath,evidence:await unchangedBaseEvidence(input.repoPath,input.pinnedBaseCommit,input.allowedScope),commands,timeoutMs:input.timeoutMs,maxOutputBytes:input.maxOutputBytes,sandbox:input.sandbox});
const checks=[...(cached?.checks??[])];
for(const c of observed.checks)if(!checks.some(x=>x.name===c.name))checks.push({name:c.name,passed:passedCheck(c),exitCode:c.exitCode,timedOut:c.timedOut});
return {baseline:{pinnedBaseCommit:input.pinnedBaseCommit,commandDigest,ranAt:new Date().toISOString(),checks},ran:commands.map(c=>c.name)};
}

/** Mark each failed check with whether it also failed on the base. A check the baseline never ran stays unmarked. */
export function markFailsOnBase(observations:ValidationCheck[],baseline:BaselineValidation):void{
for(const o of observations){if(o.passed)continue;const base=baseline.checks.find(c=>c.name===o.name);if(base)o.failsOnBase=!base.passed;}
}
/** Failed checks the Worker's change can plausibly have caused: everything that does not also fail on the base. */
export const workerCausedFailures=(observations:readonly ValidationCheck[]):ValidationCheck[]=>observations.filter(o=>!o.passed&&o.failsOnBase!==true);
const baseFailingNames=(observations:readonly ValidationCheck[]):string[]=>observations.filter(o=>!o.passed&&o.failsOnBase===true).map(o=>o.name);
/** The operator-facing stop reason when every failed check also fails on the base, so a Worker retry cannot fix any of them. */
export function baseFailureStopReason(observations:readonly ValidationCheck[]):string|undefined{
const base=baseFailingNames(observations);if(!base.length||workerCausedFailures(observations).length)return undefined;
return `Checks also fail on the base commit, so a Worker retry cannot fix them: ${base.join(', ')}. Fix the check or its network setting, then retry validation.`;
}
/** Sentence for the Orchestrator correction note naming the base-failing checks that were left out of the failure excerpt. */
export function baseFailureNote(observations:readonly ValidationCheck[]):string{
const base=baseFailingNames(observations);return base.length?`Also fails on base, not the Worker's to fix (ignore): ${base.join(', ')}. `:'';
}
Loading
Loading