feat: opt-in context filter keeps the passages of a large output that matter for the task - #141
Merged
Merged
Conversation
… matter for the task Where the context saver would keep only the head/diagnostic/tail excerpt, context.filter (beta, off by default) splits the output at line boundaries, asks Jev one score question per chunk for the agent's current task, and keeps chunks at or above minScore word for word in original order, with the final 1000 characters and marked gaps. Any failure keeps today's excerpt. The ledger, /warden status, the trace, and scripts/filter-report.mjs count filtered and excerpt outputs apart so a trial can be judged.
Turning the filter on sends whole redacted outputs to the judge and spends requests, so a project file may tune the filter but not enable it.
The ask deadline timer is unref'd, so on Node 22.19.0 the loop drained before the abort fired and the deadline test was cancelled.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
For a large tool output that no format parser fits, the context saver gives the agent a blind excerpt: a head, lines matching error/fail/warn, and a tail. Text in the middle that matters for the task is dropped, and the agent has to search the saved full output for it. Over 11 days of recorded sessions, 43 outputs took this path and the agent went back to the full output after 8 of them.
Change (beta, off by default)
context.filter(enabled: false,chunkChars2000,minScore1.5,maxKeptChars6000,timeoutMs4000). When enabled, the output is split into chunks at line boundaries, the judge scores each chunk's usefulness for the agent's current task on a four-level rubric, and chunks scoring at least 1.5 are kept word for word in original order, with the final 1000 characters always kept and gaps marked.all, duplicates, and repeated runs are unchanged. Any failure (judge error, timeout, budget, no chunk passing) keeps today's excerpt.docs/data-handling.mdand the consent text state that the whole redacted output is sent when the filter is on.context.filter.enabledis read from the user file only; a project may tune the other keys./warden statusand the trace count filtered outputs separately (recalls, kept size, requests, time, fallbacks).scripts/filter-report.mjssummarises filtered versus excerpt outputs from session files, offline.Verification
npm run check: 1222 pass / 0 fail; typecheck and build clean.tests/filter.test.tsalso passes 10 of 10 under Node 22.19.0.git log --statoutput (97 chunks) was filtered in 5 requests and 489 ms to 4,919 kept characters.