A run's cost is dominated by what the primary model reads. Work that needs no judgement — a repo-wide grep, an inventory of call sites, a log scan — costs the same per token there as the design work does. Both Claude Code and Codex can hand such a subtask to a subagent running on a smaller model, and the subagent's intermediate output never enters the primary agent's context.
CodingAgentRunner does not add an API for this. Delegation is already the CLI's own behaviour; what was missing is configuration. The runner supplies two things per run:
- Agent definitions in the place the CLI looks for them.
- A prompt-visible rule telling the primary agent that they exist and when to use them. A capability nobody is told about saves nothing.
For a Claude run, the runner writes .claude/agents/*.md into the working directory
before spawning the CLI:
| Agent | Claude model | For |
|---|---|---|
mechanical |
haiku |
Fully specified sweeps: greps, inventories, log scans, renames, extractions |
checker |
sonnet |
Mechanical verification of finished work against a stated claim |
For a Codex run the same set is written as .codex/agents/*.toml, both on
gpt-5.6-terra — the tier the Codex documentation names for lighter subagent work.
Each generated file carries the marker
generated by coding-agent-runner for this run, which makes a generated definition
recognisable in a checkout. The marker is a label, not a permission: it is not what
decides whether the runner may touch a file.
Ownership comes from creating a file, never from reading one. Definitions are
written with FileMode.CreateNew. If anything is already at
.claude/agents/mechanical.md, the create fails, and the runner leaves that file
exactly as it is — it records it in the run's inventory and moves on. This holds
whatever the file contains, so a project may commit a definition that quotes the
generated marker (a checked-out copy of one the runner produced, say) without risking
it. A repo-provided agent the runner has no default for is picked up the same way.
The runner takes back exactly what it created. When a run ends, every file this run created is deleted, and any directory it had to create is removed if it is empty again. A file or directory that already existed stays. One exception in the other direction: if a file the runner created no longer carries the marker at cleanup time, something replaced it during the run, and it is left in place rather than deleted.
The deletion happens before the run's terminal callbacks — OnFinished fires
after the workspace is already clean — so a host that commits from its completion
handler cannot pick up a generated definition. It is also guaranteed on the
exceptional paths: cleanup runs in a finally, so a terminal path that throws still
leaves the workspace as it was found.
Two consequences worth knowing. A run killed hard enough to skip its cleanup leaves
its definitions behind. The next run in that workspace will not rewrite or remove
them — it did not create them, so they are not its to touch; it adopts them as they
are and logs a warning naming the file (SubagentDefinitionLeftoverAdopted). Clearing
a stale definition is a manual step, and the marker is how you spot one. And two runs
sharing one workspace share the files: the second run adopts the first's definitions
rather than rewriting them, and only the run that created them deletes them. Give
concurrent runs their own checkout or worktree, as the clean-context model already
assumes.
Two levels, both convention rather than per-environment settings:
- One agent: commit your own
.claude/agents/<name>.md(or.codex/agents/<name>.toml). Yours wins; the runner materializes only the names you did not define. - The whole set: commit an empty
.claude/agents/.no-runner-agentsfile. The runner then materializes nothing and injects no rule for that project.
A host application can also turn it off for every run:
var runner = new CliRunner(new CliOptions
{
Delegation = new DelegationOptions { Enabled = false },
});DelegationOptions also takes a replacement agent set (Agents) and can keep the
definitions while suppressing the prompt block (InjectContextBlock = false).
The prompt the CLI receives gets the rule and the run's agent inventory appended, in the same shape as the chat-attachment block:
<delegation-economy>
Delegate simple, fully specified subtasks to one of the cheap subagents listed below
instead of doing them in this thread. ...
Subagents available for this run:
{"name":"mechanical","model":"haiku","use_for":"Mechanical, fully specified subtasks ..."}
{"name":"checker","model":"sonnet","use_for":"Mechanical verification of work that ..."}
</delegation-economy>
A project that wants to phrase the rule itself commits
contexts/delegation-economy.md in the repository root. Its text replaces the
built-in rule; the inventory is appended either way. The path is
DelegationOptions.ContextBlockRelativePath.
Delegated work is visible in the typed event stream, so a token ledger can attribute it:
- Claude emits a
tool_useframe for the delegation tool —Agentin Claude Code 2.1.220,Taskin earlier versions. The adapter maps it toToolStarted("<tool>", "<subagent name>"), reading the name from the frame'ssubagent_typerather than from the tool name, so both spellings report the agent that ran. - Frames produced inside a delegated subtask carry
parent_tool_use_id. They are otherwise ordinaryassistant/userframes and map to the usualOutputDelta/ToolCompletedevents. - While the subagent works, Claude also streams
systemframes withtask_started/task_progress/task_updated/task_notificationsubtypes. These map toHeartbeat. CliRunInfo.Subagentslists the agents a run had available, whoever provided them.
That last mapping is not cosmetic, and it is the one thing delegation broke in the
existing stream parsing. The adapter used to read every non-init system subtype as
SessionInitializing, which is right for a session frame and wrong for a subtask
reporting in, in two compounding ways. The phase fell back out of ToolExecuting
while the delegation tool was still running — swapping the watchdog's tool budget
(300 s / 1200 s) for the tighter session-initializing one (120 s / 600 s) at exactly
the point a run is doing its longest-running thing. And SessionInitializing is not
an activity signal, so task_progress — literally the subtask reporting progress —
did not reset the silence clock. A delegated sweep long enough to matter would have
been classified hung, and stopped under autoStop, while it was visibly working.
Heartbeat is the existing vocabulary for both halves: it holds the phase and counts
as liveness. The split is by prefix, so a task_* subtype added later lands on the
right side of it, and it is pinned by a test.
A demo run against a small repository (primary model claude-opus-5, prompt: list
every FIXME under src/ and tests/, then judge which matters most) delegated the
sweep and split its usage across two models:
[runner] subagents available: mechanical, checker
[text ] I'll start by finding the FIXMEs.
[tool ] Agent(mechanical)
[tool ] Bash(...) <- frames from inside the delegated subtask
[tool ] Grep(...)
...
[runner] agents dir still present after the run: False
claude-haiku-4-5-20251001 input 1016 output 2208 cache_read 38821
claude-opus-5 input 8 output 2451 cache_read 122712
The whole grep-and-inventory half of the task ran on haiku — the mechanical
agent's model — while opus kept only the judgement it was asked for.
The Codex documentation covers subagents for interactive sessions. The runner spawns
Codex through codex exec, and whether that non-interactive mode runs subagents was
not documented. Probed against codex-cli 0.145.0, three times, with and without
--enable multi_agent_v2 and -c agents.max_concurrent_threads_per_session=2:
{"type":"item.started","item":{"id":"item_1","type":"collab_tool_call","tool":"wait",
"sender_thread_id":"019fa935-...","receiver_thread_ids":[],"prompt":null,
"agents_states":{},"status":"in_progress"}}
{"type":"item.completed","item":{"id":"item_2","type":"agent_message",
"text":"Returned: MECHANICAL_OK\n\nSpawn tool used: `collaboration.spawn_agent`"}}The collab tool family is reachable in exec mode and --strict-config accepts the
agents.* configuration keys, so the config surface is real. But
receiver_thread_ids is empty and agents_states is {} — no child thread is ever
created. The model then reports a result it invented: in the probe it returned the
exact string it had been told to expect, without any agent having run.
So for Codex the runner writes the definitions (they become active the day codex exec runs subagents) but does not advertise them in the prompt. Telling the
primary agent to delegate under exec would buy fabricated results rather than cheaper
ones. That gate is SubagentSpec.AdvertiseInPrompt, and it is pinned by a test.
Delegation/SubagentDefinition.cs— one agent, as data.Delegation/SubagentDefaults.cs— the curated set.Delegation/SubagentSpec.cs— the per-CLI file convention and renderers; referenced from the CLI'sCliDescriptor, so adding a CLI is a descriptor entry.Delegation/SubagentMaterialization.cs— writing, the marker, and per-run cleanup.Delegation/DelegationContextBlock.cs— the injected block and the project override.Delegation/DelegationOptions.cs— the consumer-facing switches.