Agent Workflows runs inspectable YAML workflows. Models may generate content inside nodes; deterministic code owns dependencies, conditions, gates, retries, durable state, and evidence.
Every command supports --help. Commands that produce receipts accept
--json for one machine-readable JSON document.
piw --version
piw doctor
piw doctor --jsonshell-ready means command workflows can run. agent-ready additionally means
the optional Pi executable is available.
The default template is deterministic and zero-cost:
piw create uppercase
piw validate uppercase --strict
piw run uppercase --input hello --strict --jsonCreate a model-backed template explicitly:
piw create review --template agent \
--model openai-codex/gpt-5.6-lunaOr write steps.yaml directly:
version: 1
workflow: uppercase
input:
required: true
description: One string
steps:
- id: transform
cmd: tr '[:lower:]' '[:upper:]' < "$INPUT"
gate: tr '[:lower:]' '[:upper:]' < "$INPUT" | cmp - "$OUT"An input: block must declare both required and description—the
description is the contract a calling agent reads, so validation rejects an
input without one.
piw validate review/steps.yaml --strict
piw graph review/steps.yaml
piw graph review/steps.yaml --jsonValidation never runs a node. Strict validation rejects weak gates on nondeterministic work, including gates that only check that output exists.
piw configure review/steps.yaml work \
--model openai-codex/gpt-5.6-sol --thinking highConfigure changes only the requested node fields, preserves YAML comments, and revalidates the workflow.
piw run review/steps.yaml --input "Review this change" --strict --json
piw run review/steps.yaml --input-file request.md --strict --jsonEach run lands in runs/<workflow>-<YYYYMMDD-HHMMSS>/ beside steps.yaml.
The receipt's run field is that directory's name and is the RUN_ID
accepted by inspect and resume.
Dependency-ready nodes run concurrently—top-level workers: bounds the
pool (default 4, maximum 16). needs: is the only serialization guarantee:
nodes that must not overlap (for example commands mutating the same file) must
be ordered with explicit dependencies.
The input is copied into the run and fingerprinted. Each run freezes the workflow and records:
workflow.yamlandinput.txt—immutable execution boundary;manifest.jsonandstate.json—durable contract and current projection;trace.jsonl—contiguous committed events;ledger.json—model, time, token, and cost usage when reported;<step>.mdand<step>.stderr—artifacts and diagnostics;produced/—files declared byproduces:; and- local Git history—diffable step transitions unless explicitly disabled.
piw inspect review/steps.yaml --json
piw inspect review/steps.yaml RUN_ID --jsonDo not infer success from a model's last sentence. Check state, artifacts, gate results, trace, and ledger.
piw resume review/steps.yaml RUN_ID --jsonResume verifies the frozen workflow and immutable input and continues from the committed unfinished boundary. A changed source workflow fails closed:
piw resume review/steps.yaml RUN_ID --force-drift --jsonUse --force-drift only after reviewing the exact change. The original frozen
workflow remains in the run bundle. Input drift cannot be forced.
See the recovery example for a file-backed human checkpoint.
needs: [step-id] declares dependencies explicitly. Without needs, a node
depends on the previous listed node. References also create dependencies:
{input}—immutable run input;{prev}—previous listed step artifact;{step.id}—named prior artifact; and{run}—run directory.
Command nodes receive $INPUT, $OUT, $RUN, $STEP, and $WORKFLOW_DIR.
Use JSON output plus schema and when for deterministic decisions:
- id: decide
prompt: Return JSON with verdict pass or repair.
schema:
verdict:
type: string
enum: [pass, repair]
gate: python3 -m json.tool "$OUT" >/dev/null
- id: repair
needs: [decide]
from: decide
when:
op: equals
path: /verdict
value: repair
prompt: Repair the reported issue from {step.decide}.
gate: python3 -m json.tool "$OUT" >/dev/nullCode evaluates when; the model cannot choose which node the runner dispatches.
retries: 2
retry_on: [model_error, schema_failed, gate_failed]
retry_delay_seconds: 1
retry_backoff: exponential
retry_jitter: 0.2 # deterministic ±20% spread per step id + attempt
retry_max_delay_seconds: 60 # backoff cap (default 300)
timeout: 900Retry only declared failure classes. Commands and their child process groups are terminated when their timeout expires. Jitter is derived from the step id and attempt number, so a replayed run computes the same delays.
A step may attach an LLM judge that scores each candidate and iterates until a target score or the attempt budget is reached:
- id: draft
prompt: Write the release announcement for {input}.
gate: test -s "$OUT"
judge:
prompt: >-
Score this draft 0-10 for clarity.
Return JSON: {"score": N}. Candidate: {out}
score: 8 # minimum passing score (default 8)
max_iters: 3 # attempt budget (supersedes retries when larger)
keep_best: false # true keeps the highest-scoring rejected candidateThe judge must return JSON containing a numeric "score"; each verdict is
stored as <step>.judge<N>.md. A judge never replaces the mechanical gate—
the gate still runs on every attempt—and a below-target score is an ordinary
retryable failure class (judge_below_target). Judge and schema settings are
part of the cache key, so tightening either invalidates cached artifacts.
A top-level qa: block runs one independent model review over all artifacts
after every step has passed (and again during verification):
qa:
prompt: >-
Review these artifacts for contradictions.
Return JSON: {"verdict": "pass"} or {"verdict": "fail"}.
{artifacts}The report is written to qa.md; a fail verdict fails the run even though
every individual gate passed.
workers:(top-level, default 4, maximum 16)—concurrent dependency-ready nodes; see section 5.cwd:(top-level, default.)—execution directory for command nodes, resolved relative tosteps.yaml.preview:(step)—declarative image paths for visual tooling; never affects execution.
Use the weakest runtime that can complete the node:
cmd:—deterministic code, no model;prompt:—one isolated completion;prompt:plustools:—explicit Pi tool allowlist; orprompt:plusagent: true—full Pi tool loop.
tools: controls what Pi exposes to a model. It does not sandbox the shell,
filesystem, process, network, or inherited environment.
The complete authoring contract is
src/agent_workflows/schemas/workflow.schema.json.
Runnable examples are indexed in examples/README.md.