Skip to content

Latest commit

 

History

History
105 lines (88 loc) · 5.29 KB

File metadata and controls

105 lines (88 loc) · 5.29 KB

The agent loop

internal/agent contains one provider loop. Runtime ownership chooses its single model-facing tool (rlm_exec), while AgentSession binds that tool to the session’s kernel and host identity.

stateDiagram-v2
    [*] --> Focus: transform large input/history into handles
    Focus --> Compact: context budget exceeded
    Focus --> Model: request with rlm_exec
    Compact --> Model
    Model --> Kernel: rlm_exec call
    Kernel --> Host: Starlark module calls
    Host --> Model: bounded result/handle
    Model --> Kernel: another cell
    Model --> Commit: ordinary response
    Commit --> [*]
Loading

The ordinary assistant response completes only the current agent’s turn. It is persisted in that agent’s transcript, but it is not injected into the parent. Cross-agent communication is an explicit durable message.

Context focusing

At activation, the model receives a bounded recent history plus handles for the complete history or oversized input. context.inspect/search/read lets a cell retrieve only relevant spans. This keeps large corpora out of every model request without making them inaccessible.

Proactive compaction runs when estimated context crosses the configured fraction of the model window. A provider context-limit error may trigger one reactive compaction and retry. Compaction summaries and raw-history cutoffs are committed with the root turn.

A single provider call failing does not fail the turn. A stream that sends nothing for the stall timeout (120 s on OpenAI-compatible chat streams, 300 s on the OpenAI Responses and ChatGPT subscription streams, which can stay silent while a reasoning model thinks) is cancelled and retried, as is an attempt that hits the ten-minute per-attempt ceiling; only the caller's own deadline or cancel ends a call outright. Transport errors, 429 and 5xx responses (with Retry-After honoured up to a minute), and provider error chunks whose wording reads as transient are retried with backoff, up to maxRetries attempts before the first delta. After the first delta the partial answer cannot be resumed: within a budget of two regenerations the client discards it, tells clients through stream.discard and a notice, and requests the whole message again. Permanent failures (authentication, quota, invalid request, context limit) are never repeated. When every attempt fails, the last partial is kept in the transcript as [response interrupted] and the turn fails. Each attempt is admitted and settled separately in model accounting.

A fold keeps the system prompt, one running summary, and a token-budgeted tail of recent whole turns. When the newest turn alone exceeds the tail budget, as a long tool-heavy turn does, the tail boundary moves inside that turn onto an assistant/tool-pair boundary and the turn's opening user message is kept verbatim between the summary and the tail, so the model keeps acting on its exact instructions. Before this rule existed the newest turn was always kept whole, so a single turn of dozens of tool exchanges could never be folded; each round re-summarized only the prior summary, and one agent spent 40 minutes and 52 model calls making six tool calls of progress (.ai-docs/plans/compaction-loop). Two guards now stop that loop: a fold with nothing left to fold makes no model call, and a fold that cannot get back under the threshold stalls further proactive folds for the rest of the turn. The window itself is left to the provider: a rejection with nothing left to fold fails the turn with ErrCompactionExhausted, which reaches the parent through the usual failure notice. While a turn runs, last_turn on the agent record carries model_calls, compactions, and last_activity_at, so a parent polling agents.list or a client rendering the agent row can tell steady work from a stalled loop. On reload the pinned message is re-derived from the raw log, so a resumed agent sees the same view the live one had.

Child activation

agents.spawn performs these steps:

  1. validate requested capabilities, budgets, and ancestry;
  2. build an identical AgentSession and reserve a kernel worker;
  3. commit the retained child and delegated grants atomically;
  4. launch its first turn asynchronously.

Later message or agent-change notifications can activate the retained child again. Notifications are coalesced metadata; the child inspects durable state to decide what to do. On restart, the daemon reconstructs retained sessions, their focused transcripts, authority, model route, and kernels.

The default recursion limit is two edges: root → child → grandchild.

Model fan-out

models.batch runs independent completion calls concurrently and returns results in input order. These calls share the caller’s durable token, cost, elapsed, and active-operation budgets but do not create agent identities.

Failure behavior

  • A failed cell returns an error to the tool loop; committed host state is not rolled back speculatively.
  • A worker crash loses Starlark globals only. The next execution starts a new worker against durable host state.
  • A child turn may fail and remain retained/idle for a later activation.
  • Stopping or deleting a child terminalizes its whole subtree and cancels its live processes and kernels.
  • Daemon recovery keeps committed outcomes and marks uncertain work interrupted rather than replaying side effects.