Repository navigation
feat(arcan-core): BRO-425 context compression middleware — SummarizationMiddleware - #1786
Conversation
…sion (BRO-425) Add automatic context compression for long agent sessions. When estimated context exceeds a configurable token threshold, older conversation turns are folded into a compressed summary while the most recent N turns are kept at full fidelity. Full history is never lost — compression rewrites only the per-call ProviderRequest (a clone of the durable message log), so the Lago event journal upstream stays complete and replayable. - context_compiler: new `ContextBlockKind::Compressed` variant (assembly order after Memory, before Retrieval) + default per-kind budget + `compressed_block` helper to build the typed block for the system-prompt compiler. - summarization: `SummarizationMiddleware` (impl `TurnMiddleware`), pluggable `Summarizer` trait with a deterministic LLM-free `HeuristicSummarizer` default (an LLM-backed summarizer drops in behind the trait, fail-open by design), and `SummarizationConfig` exposing all three thresholds (token limit, recent turn count, per-message result-size cap). - Turns are user-message-delimited; system messages always preserved verbatim; a "won't shrink" guard skips rewriting when a summary wouldn't reduce tokens. Tests (arcan-core, all green): recent-N verbatim preservation, key-decision survival across a long many-tool-call session, durable-history-untouched via a recording provider through Orchestrator::with_turn_middlewares, custom summarizer dispatch, threshold/no-op boundaries, and Compressed-block assembly order. Scope note: consumes the arcan_core::TurnMiddleware chain (Orchestrator). The production KernelRuntime uses aios-runtime's onion-style TurnMiddleware — wiring this compressor there needs a TurnContext-level adapter and is left as a follow-up. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
📝 WalkthroughWalkthroughAdds automatic summarization middleware for long sessions, introduces the ChangesContext compression
Estimated code review effort: 4 (Complex) | ~45 minutes Sequence Diagram(s)sequenceDiagram
participant Runtime as Runtime orchestrator
participant Middleware as SummarizationMiddleware
participant Summarizer as HeuristicSummarizer
participant Provider as Model provider
Runtime->>Middleware: before_model_call(ProviderRequest)
Middleware->>Middleware: estimate tokens and split turns
Middleware->>Summarizer: summarize older turns
Summarizer-->>Middleware: compressed summary message
Middleware->>Provider: send rewritten messages
Provider-->>Runtime: model response
Runtime->>Runtime: retain complete output history
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
crates/arcan/arcan-core/src/context_compiler.rs (1)
62-77: 🎯 Functional Correctness | 🟠 Major | ⚡ Quick winDefault per-kind budgets now exceed
total_budget, breaking the previous exact-fit invariant.Before this change, the default
block_budgetssummed to exactly30_000(matchingtotal_budget), so a fully-populated context (each block at its per-kind cap) never triggered the drop-lowest-priority path. Adding(Compressed, 6_000)brings the sum to36_000whiletotal_budgetstays30_000. Under default config, a session using all seven block kinds near their caps will now silently drop lower-priority blocks (e.g.Task,Workspace) even though each individual block respects its own budget — an unintended regression with no test coverage for the sum invariant.🔧 Proposed fix: raise `total_budget` to match the new budget sum (or shrink another kind's budget accordingly)
impl Default for ContextCompilerConfig { fn default() -> Self { Self { - total_budget: 30_000, + total_budget: 36_000, block_budgets: vec![ (ContextBlockKind::Persona, 2_000), (ContextBlockKind::Rules, 5_000), (ContextBlockKind::Memory, 8_000), (ContextBlockKind::Compressed, 6_000), (ContextBlockKind::Retrieval, 6_000), (ContextBlockKind::Workspace, 5_000), (ContextBlockKind::Task, 4_000), ], } } }🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@crates/arcan/arcan-core/src/context_compiler.rs` around lines 62 - 77, Update the default values in ContextCompilerConfig::default so total_budget matches the sum of all entries in block_budgets, preserving the exact-fit invariant with the newly added Compressed budget. Keep each per-kind budget unchanged unless necessary to maintain that invariant.
🧹 Nitpick comments (1)
crates/arcan/arcan-core/src/summarization.rs (1)
207-307: 🚀 Performance & Scalability | 🔵 Trivial | 🏗️ Heavy liftConsider caching/incrementally reusing the summary to avoid quadratic cost over very long sessions.
compress()re-derivesolder/recentand re-runsSummarizer::summarizeover the entire older-turn history on everybefore_model_call, from scratch each time (nothing is cached across calls). For the module's stated target — long tool-driven sessions with many iterations — total work across a run grows roughly with the square of the conversation length, since each of the O(n) iterations re-summarizes an O(n)-sized (and growing) older-turn slice. This is functionally correct (fail-open, bounded per call) but could become a real cost center for the very sessions this middleware is meant to help.Worth considering for a follow-up: persist/cache the last computed summary (and its turn boundary) on the middleware or in session state, and only re-summarize the newly-aged turns since the last compression pass, rather than the full older window every time.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@crates/arcan/arcan-core/src/summarization.rs` around lines 207 - 307, Update SummarizationMiddleware::compress to reuse summary work across repeated calls by caching the previously summarized turn boundary and summary, preferably in appropriate middleware or session state. When additional turns age beyond the preserved recent window, summarize only those newly eligible turns and combine them with the cached summary; preserve existing ordering, token-threshold behavior, fail-open behavior, and CompressionOutcome semantics.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Outside diff comments:
In `@crates/arcan/arcan-core/src/context_compiler.rs`:
- Around line 62-77: Update the default values in ContextCompilerConfig::default
so total_budget matches the sum of all entries in block_budgets, preserving the
exact-fit invariant with the newly added Compressed budget. Keep each per-kind
budget unchanged unless necessary to maintain that invariant.
---
Nitpick comments:
In `@crates/arcan/arcan-core/src/summarization.rs`:
- Around line 207-307: Update SummarizationMiddleware::compress to reuse summary
work across repeated calls by caching the previously summarized turn boundary
and summary, preferably in appropriate middleware or session state. When
additional turns age beyond the preserved recent window, summarize only those
newly eligible turns and combine them with the cached summary; preserve existing
ordering, token-threshold behavior, fail-open behavior, and CompressionOutcome
semantics.
ℹ️ Review info
⚙️ Run configuration
Configuration used: defaults
Review profile: CHILL
Plan: Pro Plus
Run ID: 68b49c1d-4bc7-4961-b2ce-a7323b1a896b
📒 Files selected for processing (4)
crates/arcan/arcan-core/src/context_compiler.rscrates/arcan/arcan-core/src/lib.rscrates/arcan/arcan-core/src/runtime.rscrates/arcan/arcan-core/src/summarization.rs
P20 Cross-Model Adversarial Review (Strata B — fresh-context devil's-advocate)SCORE: 8/10 — VERDICT: REAL/CORRECT and mergeable (≥7 threshold met → merge authorized) The reviewer treated the diff as untrusted data and attempted to refute correctness + slop-freedom across correctness / slop / claims-vs-reality / integration-risk. Could not refute correctness: turn splitting (user-delimited), Held below 9 (non-blocking):
Audit trail per P20; proceeding to p9 auto-merge lifecycle. |
BRO-425 — Context Compression Middleware
Adds automatic context compression for long agent sessions. When a session's
estimated context exceeds a configurable token threshold, older conversation
turns are folded into a compressed summary while the most recent N turns stay
at full fidelity — mirroring DeerFlow's
SummarizationMiddlewareand Hermes'context_compressor.py, adapted to Life's typed context compiler.What changed (target crate:
arcan-core)context_compiler.rsContextBlockKind::Compressedvariant (assembly order: afterMemory,before
Retrieval— groups "what happened before" ahead of task blocks).default_config_reasonableupdated (6→7).compressed_block(summary, priority)helper builds the typed block for callersassembling the system prompt via
compile_context.summarization.rs(new)SummarizationMiddlewareimplementingarcan_core::TurnMiddleware. Inbefore_model_callit estimates context tokens and, over threshold, keeps allsystem messages + the most recent N turns verbatim and folds older turns into a
single compressed summary system message.
Summarizertrait (pluggable) + deterministic, LLM-freeHeuristicSummarizerdefault. An LLM-backed summarizer drops in behind the trait;
summarizeisinfallible so the middleware is fail-open (never blocks a model call).
SummarizationConfigexposes all three thresholds:token_threshold,recent_turns(default 10),result_char_threshold.a summary wouldn't actually reduce tokens.
Full history is never lost
Compression rewrites only the per-call
ProviderRequest.messages, which is aclone of the orchestrator's durable message log. The canonical history — and
therefore the Lago event journal upstream — is untouched and stays replayable via
lago log/lago replay. A test assertsoutput.messages.len() > seen.len()(durable history uncompressed while the model saw the compressed copy).
Acceptance criteria
CompressedBlockvariant in context compiler (ContextBlockKind::Compressed)Validation
cargo test -p arcan-core→ 160 lib tests + 1 doc test, all green(new: recent-N verbatim preservation, key-decision survival across a 40-tool-call
session, durable-history-untouched via a recording provider through
Orchestrator::with_turn_middlewares, custom-summarizer dispatch, threshold/no-opboundaries,
Compressed-block assembly order).cargo clippy -p arcan-core --all-targets→ clean.cargo checkon downstream consumers (arcan-lago,arcan-fleet,aios-runtime,nous-middleware) → compile clean (new enum variant breaks no exhaustive match).Scope note
The middleware consumes the
arcan_core::TurnMiddlewarechain (used byarcan_core::Orchestrator, where it's tested end-to-end). The productionKernelRuntimeusesaios-runtime's onion-styleTurnMiddleware(a distincttrait over
TurnContext); wiring this compressor into that path needs aTurnContext-level adapter and is intentionally left as a follow-up — kept out tokeep this change surgical and within the ticket's stated target crate.
🤖 Generated with Claude Code
Summary by CodeRabbit
New Features
Tests