internal/agent contains one provider loop. Runtime ownership chooses its
single model-facing tool (rlm_exec), while AgentSession binds that tool to
the session’s kernel and host identity.
stateDiagram-v2
[*] --> Focus: transform large input/history into handles
Focus --> Compact: context budget exceeded
Focus --> Model: request with rlm_exec
Compact --> Model
Model --> Kernel: rlm_exec call
Kernel --> Host: Starlark module calls
Host --> Model: bounded result/handle
Model --> Kernel: another cell
Model --> Commit: ordinary response
Commit --> [*]
The ordinary assistant response completes only the current agent’s turn. It is persisted in that agent’s transcript, but it is not injected into the parent. Cross-agent communication is an explicit durable message.
At activation, the model receives a bounded recent history plus handles for
the complete history or oversized input. context.inspect/search/read lets a
cell retrieve only relevant spans. This keeps large corpora out of every model
request without making them inaccessible.
Proactive compaction runs when estimated context crosses the configured fraction of the model window. A provider context-limit error may trigger one reactive compaction and retry. Compaction summaries and raw-history cutoffs are committed with the root turn.
A single provider call failing does not fail the turn. A stream that sends
nothing for the stall timeout (120 s on OpenAI-compatible chat streams, 300 s
on the OpenAI Responses and ChatGPT subscription streams, which can stay silent
while a reasoning model thinks) is cancelled and retried, as is an attempt that
hits the ten-minute per-attempt ceiling; only the caller's own deadline or
cancel ends a call outright. Transport errors, 429 and 5xx responses (with
Retry-After honoured up to a minute), and provider error chunks whose wording
reads as transient are retried with backoff, up to maxRetries attempts before
the first delta. After the first delta the partial answer cannot be resumed:
within a budget of two regenerations the client discards it, tells clients
through stream.discard and a notice, and requests the whole message again.
Permanent failures (authentication, quota, invalid request, context limit) are
never repeated. When every attempt fails, the last partial is kept in the
transcript as [response interrupted] and the turn fails. Each attempt is
admitted and settled separately in model accounting.
A fold keeps the system prompt, one running summary, and a token-budgeted
tail of recent whole turns. When the newest turn alone exceeds the tail
budget, as a long tool-heavy turn does, the tail boundary moves inside that
turn onto an assistant/tool-pair boundary and the turn's opening user message
is kept verbatim between the summary and the tail, so the model keeps acting
on its exact instructions. Before this rule existed the newest turn was always
kept whole, so a single turn of dozens of tool exchanges could never be
folded; each round re-summarized only the prior summary, and one agent spent
40 minutes and 52 model calls making six tool calls of progress
(.ai-docs/plans/compaction-loop). Two guards now stop that loop: a fold with
nothing left to fold makes no model call, and a fold that cannot get back
under the threshold stalls further proactive folds for the rest of the turn.
The window itself is left to the provider: a rejection with nothing left to
fold fails the turn with ErrCompactionExhausted, which reaches the parent
through the usual failure notice. While a turn runs, last_turn on the agent
record carries model_calls, compactions, and last_activity_at, so a
parent polling agents.list or a client rendering the agent row can tell
steady work from a stalled loop. On reload the pinned message is re-derived
from the raw log, so a resumed agent sees the same view the live one had.
agents.spawn performs these steps:
- validate requested capabilities, budgets, and ancestry;
- build an identical
AgentSessionand reserve a kernel worker; - commit the retained child and delegated grants atomically;
- launch its first turn asynchronously.
Later message or agent-change notifications can activate the retained child again. Notifications are coalesced metadata; the child inspects durable state to decide what to do. On restart, the daemon reconstructs retained sessions, their focused transcripts, authority, model route, and kernels.
The default recursion limit is two edges: root → child → grandchild.
models.batch runs independent completion calls concurrently and returns
results in input order. These calls share the caller’s durable token, cost,
elapsed, and active-operation budgets but do not create agent identities.
- A failed cell returns an error to the tool loop; committed host state is not rolled back speculatively.
- A worker crash loses Starlark globals only. The next execution starts a new worker against durable host state.
- A child turn may fail and remain retained/idle for a later activation.
- Stopping or deleting a child terminalizes its whole subtree and cancels its live processes and kernels.
- Daemon recovery keeps committed outcomes and marks uncertain work interrupted rather than replaying side effects.