Problem
When a tool prints more than context.tailMinChars (default 12000), the context saver stores the full output in a file and replaces it in the agent's context with an excerpt. Today the excerpt is one of two things:
- a format-specific parse when Jev names a format with confidence above
context.formatConfidence (test runner output, a diff, a stack trace; formatExcerpt in src/excerpt.ts), or
- a generic head and tail (
compressOutput in src/output.ts).
Neither knows what the agent was looking for. A 40 KB npm test run where the one relevant assertion failure sits in the middle loses it to head-and-tail. The agent then recalls the full file, which undoes the saving and costs a round trip. The recall ledger (src/recall.ts) measures how often this happens.
Proposal
Task-aware segment selection, as a third excerpt strategy ahead of the generic one.
- Split deterministically in code. Cut the output into at most 16 segments on natural boundaries: blank-line runs, test-case headers,
diff --git markers, stack frames, log timestamps, or fixed-size fallback when no boundary is found. Each segment gets an index and a one-line label (first non-blank line, clipped).
- One request, one Noul per segment. In the same output request the saver already sends, add
segment_<i> questions of the form "Does segment i contain what the agent needs for the task, given task and the tool's command?" The task is the latest user turn, already in the request state. Sixteen segments plus the existing retention and format questions stay under the 32-question cap.
- Keep the selected segments verbatim. Segments at or above
context.segmentKeep (new key, default around 0.7) are kept in full, in original order, with a one-line marker where dropped segments were: [segment 4 of 16 omitted, 2.1 KB, "PASS src/auth/login.test.ts"]. The full output path and the scoped-recall instruction stay at the end as today.
- Bound the result. If the kept segments exceed a size ceiling (say 40 percent of the original or a fixed character cap), keep the highest-scoring ones until the ceiling is met.
- Fall back. With no segment above the threshold, or a request failure, use the existing format or head-and-tail path. Nothing gets worse.
Calibration
scripts/context-cases.mjs holds labelled synthetic outputs for the retention and format questions. Extend each case with expectedSegments: number[] and score precision and recall of the selection. Tune segmentKeep on that corpus. Recall rate from the ledger, before and after, is the live metric: fewer whole-file recalls means the excerpt kept what mattered.
Open questions for discussion
- Segment boundary rules per format: reuse the format parsers' structure where a format is detected, generic boundaries otherwise?
- Should the question see the segment text or only its label plus size? Text is more accurate and costs more bytes per request; the request already carries the output head and tail.
- Should selection also apply to duplicate-output notes, or only to first-time large outputs?
Non-goals
- Summarising segments with a model. Kept segments are verbatim; dropped segments are named, not paraphrased.
- Changing the recall tool or the saved-file format.
Problem
When a tool prints more than
context.tailMinChars(default 12000), the context saver stores the full output in a file and replaces it in the agent's context with an excerpt. Today the excerpt is one of two things:context.formatConfidence(test runner output, a diff, a stack trace;formatExcerptinsrc/excerpt.ts), orcompressOutputinsrc/output.ts).Neither knows what the agent was looking for. A 40 KB
npm testrun where the one relevant assertion failure sits in the middle loses it to head-and-tail. The agent then recalls the full file, which undoes the saving and costs a round trip. The recall ledger (src/recall.ts) measures how often this happens.Proposal
Task-aware segment selection, as a third excerpt strategy ahead of the generic one.
diff --gitmarkers, stack frames, log timestamps, or fixed-size fallback when no boundary is found. Each segment gets an index and a one-line label (first non-blank line, clipped).segment_<i>questions of the form "Does segment i contain what the agent needs for the task, giventaskand the tool's command?" The task is the latest user turn, already in the request state. Sixteen segments plus the existing retention and format questions stay under the 32-question cap.context.segmentKeep(new key, default around 0.7) are kept in full, in original order, with a one-line marker where dropped segments were:[segment 4 of 16 omitted, 2.1 KB, "PASS src/auth/login.test.ts"]. The full output path and the scoped-recall instruction stay at the end as today.Calibration
scripts/context-cases.mjsholds labelled synthetic outputs for the retention and format questions. Extend each case withexpectedSegments: number[]and score precision and recall of the selection. TunesegmentKeepon that corpus. Recall rate from the ledger, before and after, is the live metric: fewer whole-file recalls means the excerpt kept what mattered.Open questions for discussion
Non-goals