Skip to content

Context saver: recall-driven per-family threshold calibration #65

Description

@DevMortimer

Problem

The context saver decides per output whether to compress and how much to keep (Retention in src/output.ts: all, errors_and_summary, summary_only). The threshold is one global context.confidence (default 0.8). The session ledger (src/saver.ts, src/recall.ts) records, for each compressed output, whether the agent later recalled the stored file, and whether the recall was whole-file or scoped.

A whole-file recall means the excerpt was not enough for that command family: the saving was undone and a round trip was spent. Today that signal is shown in /warden status and nothing acts on it. Some families compress well (git log, ls -R, find), others do not (npm test on a project whose failures land mid-output, cargo build with errors after warnings). One global threshold cannot fit both.

Proposal

Per command family, raise the keep threshold when the family's recall rate is high.

  1. Persist samples. Append { family, tool, retention, format, recalled, recallKind, ts } to a project-local ledger file (for example .local/warden-context.jsonl, gitignored, redacted, no output text). The in-memory ledger already has all fields but the family; derive the family the same way the action guard and stuck detector do so all three agree.
  2. Adjust per family. At session start, load the ledger. For a family with at least 5 samples and a whole-file recall rate above 10 percent, raise the effective threshold for that family by a step (say 0.05 per 10 points of recall, capped at 0.95). A family with 20 samples and under 3 percent recall may lower it by one step toward the global default. Scoped recalls do not count against a family; they are the intended path.
  3. Show it. /warden status lists families with an adjusted threshold and their sample count and recall rate. A context.adaptive: false key turns adjustment off and uses the global threshold everywhere.
  4. Bound it. The adjustment never goes below the global default minus one step or above 0.95. A family's history is capped at the newest 200 samples.

Acceptance

  • Unit tests on a pure function adjustedThreshold(global, samples) covering: fewer than 5 samples returns global; 10 percent recall over 5 samples raises by one step; 30 percent raises by three; cap at 0.95; scoped recalls ignored; 20 samples at 2 percent lowers by one step.
  • Ledger read and write tests with a temp directory; a corrupt line is skipped, not fatal.
  • Status output test.
  • npm run check passes. Docs for the key and the file in docs/configuration.md and docs/data-handling.md (the file stays local). CHANGELOG entry under Unreleased.

Open questions for discussion

  • Family granularity: npm test and npm run check as one family or two? The stuck detector's family function is the candidate to reuse; confirm it splits where recall behaviour differs.
  • Should the ledger be per project or per user? Per project matches the config override model; per user learns faster.
  • Interaction with task-aware segment selection, if that lands first: recall rate then measures the selector, and the adjustment should tune segmentKeep rather than the retention threshold.

Non-goals

  • Learning from anything other than the agent's own recalls.
  • Sending recall data anywhere.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requestquestionFurther information is requested

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions