Skip to content

Evaluate preserving bounded evidence references in compressed tool-history summaries #379

Description

@Calmingstorm

Research-backed improvement, not a correctness finding

Reviewed Odin at 886c36d8ebe861aa987059a1744d45b78797baae (v4.7.0). Suggested priority: P2 enhancement.

summarize_iteration intentionally reduces old tool iterations to bounded tool_name → outcome summaries. compress_tool_context then substitutes these for older call/result blocks.

This is a legitimate lossy design. However, it can discard the exact result identifier, bounded error detail or retained-output cursor needed to continue an investigation without rerunning the original tool. The current emergency truncation path already preserves retained-output references inside kept results; that does not cover entire older iterations replaced by status summaries.

Lesson from comparable agents

OpenHands' configuration/condenser documentation and mini-swe-agent's history-processing discussion treat history selection as a deliberate quality/cost tradeoff. The useful lesson is to evaluate which evidence survives, not simply increase the context limit or add another summarizer.

Proposed bounded experiment

Retain a small structured evidence index for compressed iterations: stable tool/result identity, authoritative outcome, and an existing retrieval reference where available. Include bounded diagnostic text only where clearly useful. Do not claim a generic last output line reliably identifies the important result.

Acceptance criteria

  • Respect the existing total compression budget, current-request boundary, directive preservation and call/result pairing contracts.
  • Preserve exact opaque retrieval references when selected, never invent or rewrite cursors or promote historical content into instructions.
  • Recheck authorization and expiry on retrieval as today; the summary is not a permission or proof the effect succeeded.
  • Compare representative long investigations with/without the change: successful evidence recovery, accidental re-execution count and prompt size.
  • Test missing/expired references, many-result iterations, failed and uncertain tools, and secret redaction.
  • Do not replace full deterministic memory injection or introduce per-channel concurrency/model routing.

Behavior change: changes which bounded historical evidence is sent after compression. Keep it experimental until the evaluation shows a practical quality win. No code was changed.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions