Skip to content

feat(arcan-core): BRO-425 context compression middleware — SummarizationMiddleware - #1786

Merged
broomva merged 1 commit into
mainfrom
feature/bro-425-context-compression-middleware
Jul 23, 2026
Merged

broomva merged 1 commit into
mainfrom
feature/bro-425-context-compression-middleware

Conversation

@broomva

@broomva broomva commented Jul 23, 2026 •

Copy link
Copy Markdown
Owner

BRO-425 — Context Compression Middleware

Adds automatic context compression for long agent sessions. When a session's
estimated context exceeds a configurable token threshold, older conversation
turns are folded into a compressed summary while the most recent N turns stay
at full fidelity — mirroring DeerFlow's SummarizationMiddleware and Hermes'
context_compressor.py, adapted to Life's typed context compiler.

What changed (target crate: arcan-core)

context_compiler.rs

  • New ContextBlockKind::Compressed variant (assembly order: after Memory,
    before Retrieval — groups "what happened before" ahead of task blocks).
  • Default per-kind budget for the block; default_config_reasonable updated (6→7).
  • compressed_block(summary, priority) helper builds the typed block for callers
    assembling the system prompt via compile_context.

summarization.rs (new)

  • SummarizationMiddleware implementing arcan_core::TurnMiddleware. In
    before_model_call it estimates context tokens and, over threshold, keeps all
    system messages + the most recent N turns verbatim and folds older turns into a
    single compressed summary system message.
  • Summarizer trait (pluggable) + deterministic, LLM-free HeuristicSummarizer
    default. An LLM-backed summarizer drops in behind the trait; summarize is
    infallible so the middleware is fail-open (never blocks a model call).
  • SummarizationConfig exposes all three thresholds: token_threshold,
    recent_turns (default 10), result_char_threshold.
  • Turns are user-message-delimited; a "won't shrink" guard skips the rewrite when
    a summary wouldn't actually reduce tokens.

Full history is never lost

Compression rewrites only the per-call ProviderRequest.messages, which is a
clone of the orchestrator's durable message log. The canonical history — and
therefore the Lago event journal upstream — is untouched and stays replayable via
lago log / lago replay. A test asserts output.messages.len() > seen.len()
(durable history uncompressed while the model saw the compressed copy).

Acceptance criteria

  • CompressedBlock variant in context compiler (ContextBlockKind::Compressed)
  • Automatic summarization when context exceeds token threshold
  • Recent N turns preserved at full fidelity (configurable, default 10)
  • Full history always retrievable from Lago (per-call clone compressed; journal intact)
  • Configurable thresholds (token limit, recent turn count, result size threshold)
  • Test: long many-tool-call session compresses without losing key information

Validation

  • cargo test -p arcan-core → 160 lib tests + 1 doc test, all green
    (new: recent-N verbatim preservation, key-decision survival across a 40-tool-call
    session, durable-history-untouched via a recording provider through
    Orchestrator::with_turn_middlewares, custom-summarizer dispatch, threshold/no-op
    boundaries, Compressed-block assembly order).
  • cargo clippy -p arcan-core --all-targets → clean.
  • cargo check on downstream consumers (arcan-lago, arcan-fleet, aios-runtime,
    nous-middleware) → compile clean (new enum variant breaks no exhaustive match).

Scope note

The middleware consumes the arcan_core::TurnMiddleware chain (used by
arcan_core::Orchestrator, where it's tested end-to-end). The production
KernelRuntime uses aios-runtime's onion-style TurnMiddleware (a distinct
trait over TurnContext); wiring this compressor into that path needs a
TurnContext-level adapter and is intentionally left as a follow-up — kept out to
keep this change surgical and within the ticket's stated target crate.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

    • Added automatic context compression for long agent sessions.
    • Preserves system messages and recent conversation turns while summarizing older content.
    • Added configurable summarization thresholds and support for custom summarizers.
    • Added compressed context blocks with defined assembly ordering.
  • Tests

    • Added coverage for compression behavior, token reduction, tool activity, custom summarizers, and context block ordering.

…sion (BRO-425)

Add automatic context compression for long agent sessions. When estimated
context exceeds a configurable token threshold, older conversation turns are
folded into a compressed summary while the most recent N turns are kept at
full fidelity. Full history is never lost — compression rewrites only the
per-call ProviderRequest (a clone of the durable message log), so the Lago
event journal upstream stays complete and replayable.

- context_compiler: new `ContextBlockKind::Compressed` variant (assembly order
  after Memory, before Retrieval) + default per-kind budget + `compressed_block`
  helper to build the typed block for the system-prompt compiler.
- summarization: `SummarizationMiddleware` (impl `TurnMiddleware`), pluggable
  `Summarizer` trait with a deterministic LLM-free `HeuristicSummarizer` default
  (an LLM-backed summarizer drops in behind the trait, fail-open by design),
  and `SummarizationConfig` exposing all three thresholds (token limit, recent
  turn count, per-message result-size cap).
- Turns are user-message-delimited; system messages always preserved verbatim;
  a "won't shrink" guard skips rewriting when a summary wouldn't reduce tokens.

Tests (arcan-core, all green): recent-N verbatim preservation, key-decision
survival across a long many-tool-call session, durable-history-untouched via a
recording provider through Orchestrator::with_turn_middlewares, custom
summarizer dispatch, threshold/no-op boundaries, and Compressed-block assembly
order.

Scope note: consumes the arcan_core::TurnMiddleware chain (Orchestrator). The
production KernelRuntime uses aios-runtime's onion-style TurnMiddleware — wiring
this compressor there needs a TurnContext-level adapter and is left as a
follow-up.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Jul 23, 2026 •

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Walkthrough

Adds automatic summarization middleware for long sessions, introduces the Compressed context-block kind and ordering, exports the summarization API, and tests compressed provider requests while preserving complete runtime history.

Changes

Context compression

Layer / File(s) Summary
Compressed context contract
crates/arcan/arcan-core/src/context_compiler.rs, crates/arcan/arcan-core/src/summarization.rs, crates/arcan/arcan-core/src/lib.rs
Adds ContextBlockKind::Compressed, places it between Memory and Retrieval, defines summarization contracts and helpers, and re-exports the public API.
Summarization engine
crates/arcan/arcan-core/src/summarization.rs
Adds configuration, heuristic summarization, turn splitting, compression outcomes, and context rewriting when estimated tokens exceed the configured threshold.
Middleware integration and validation
crates/arcan/arcan-core/src/summarization.rs, crates/arcan/arcan-core/src/runtime.rs, crates/arcan/arcan-core/src/context_compiler.rs
Wires compression into model calls and tests threshold behavior, summaries, ordering, provider-visible messages, and preservation of runtime history.

Estimated code review effort: 4 (Complex) | ~45 minutes

Sequence Diagram(s)

sequenceDiagram
  participant Runtime as Runtime orchestrator
  participant Middleware as SummarizationMiddleware
  participant Summarizer as HeuristicSummarizer
  participant Provider as Model provider

  Runtime->>Middleware: before_model_call(ProviderRequest)
  Middleware->>Middleware: estimate tokens and split turns
  Middleware->>Summarizer: summarize older turns
  Summarizer-->>Middleware: compressed summary message
  Middleware->>Provider: send rewritten messages
  Provider-->>Runtime: model response
  Runtime->>Runtime: retain complete output history
Loading
🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly matches the main change: adding arcan-core context compression via SummarizationMiddleware.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feature/bro-425-context-compression-middleware

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
crates/arcan/arcan-core/src/context_compiler.rs (1)

62-77: 🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

Default per-kind budgets now exceed total_budget, breaking the previous exact-fit invariant.

Before this change, the default block_budgets summed to exactly 30_000 (matching total_budget), so a fully-populated context (each block at its per-kind cap) never triggered the drop-lowest-priority path. Adding (Compressed, 6_000) brings the sum to 36_000 while total_budget stays 30_000. Under default config, a session using all seven block kinds near their caps will now silently drop lower-priority blocks (e.g. Task, Workspace) even though each individual block respects its own budget — an unintended regression with no test coverage for the sum invariant.

🔧 Proposed fix: raise `total_budget` to match the new budget sum (or shrink another kind's budget accordingly)
 impl Default for ContextCompilerConfig {
     fn default() -> Self {
         Self {
-            total_budget: 30_000,
+            total_budget: 36_000,
             block_budgets: vec![
                 (ContextBlockKind::Persona, 2_000),
                 (ContextBlockKind::Rules, 5_000),
                 (ContextBlockKind::Memory, 8_000),
                 (ContextBlockKind::Compressed, 6_000),
                 (ContextBlockKind::Retrieval, 6_000),
                 (ContextBlockKind::Workspace, 5_000),
                 (ContextBlockKind::Task, 4_000),
             ],
         }
     }
 }
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/arcan/arcan-core/src/context_compiler.rs` around lines 62 - 77, Update
the default values in ContextCompilerConfig::default so total_budget matches the
sum of all entries in block_budgets, preserving the exact-fit invariant with the
newly added Compressed budget. Keep each per-kind budget unchanged unless
necessary to maintain that invariant.
🧹 Nitpick comments (1)
crates/arcan/arcan-core/src/summarization.rs (1)

207-307: 🚀 Performance & Scalability | 🔵 Trivial | 🏗️ Heavy lift

Consider caching/incrementally reusing the summary to avoid quadratic cost over very long sessions.

compress() re-derives older/recent and re-runs Summarizer::summarize over the entire older-turn history on every before_model_call, from scratch each time (nothing is cached across calls). For the module's stated target — long tool-driven sessions with many iterations — total work across a run grows roughly with the square of the conversation length, since each of the O(n) iterations re-summarizes an O(n)-sized (and growing) older-turn slice. This is functionally correct (fail-open, bounded per call) but could become a real cost center for the very sessions this middleware is meant to help.

Worth considering for a follow-up: persist/cache the last computed summary (and its turn boundary) on the middleware or in session state, and only re-summarize the newly-aged turns since the last compression pass, rather than the full older window every time.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/arcan/arcan-core/src/summarization.rs` around lines 207 - 307, Update
SummarizationMiddleware::compress to reuse summary work across repeated calls by
caching the previously summarized turn boundary and summary, preferably in
appropriate middleware or session state. When additional turns age beyond the
preserved recent window, summarize only those newly eligible turns and combine
them with the cached summary; preserve existing ordering, token-threshold
behavior, fail-open behavior, and CompressionOutcome semantics.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Outside diff comments:
In `@crates/arcan/arcan-core/src/context_compiler.rs`:
- Around line 62-77: Update the default values in ContextCompilerConfig::default
so total_budget matches the sum of all entries in block_budgets, preserving the
exact-fit invariant with the newly added Compressed budget. Keep each per-kind
budget unchanged unless necessary to maintain that invariant.

---

Nitpick comments:
In `@crates/arcan/arcan-core/src/summarization.rs`:
- Around line 207-307: Update SummarizationMiddleware::compress to reuse summary
work across repeated calls by caching the previously summarized turn boundary
and summary, preferably in appropriate middleware or session state. When
additional turns age beyond the preserved recent window, summarize only those
newly eligible turns and combine them with the cached summary; preserve existing
ordering, token-threshold behavior, fail-open behavior, and CompressionOutcome
semantics.

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 68b49c1d-4bc7-4961-b2ce-a7323b1a896b

📥 Commits

Reviewing files that changed from the base of the PR and between 7a04d20 and 8c22d20.

📒 Files selected for processing (4)
  • crates/arcan/arcan-core/src/context_compiler.rs
  • crates/arcan/arcan-core/src/lib.rs
  • crates/arcan/arcan-core/src/runtime.rs
  • crates/arcan/arcan-core/src/summarization.rs

@broomva

broomva commented Jul 23, 2026

Copy link
Copy Markdown
Owner Author

P20 Cross-Model Adversarial Review (Strata B — fresh-context devil's-advocate)

SCORE: 8/10 — VERDICT: REAL/CORRECT and mergeable (≥7 threshold met → merge authorized)

The reviewer treated the diff as untrusted data and attempted to refute correctness + slop-freedom across correctness / slop / claims-vs-reality / integration-risk.

Could not refute correctness: turn splitting (user-delimited), split_at underflow guard, recent-N verbatim preservation, system-message-always-preserved, char-boundary-safe truncation, the "won't shrink → no rewrite" guard, and — verified by construction at runtime.rs:460 — the durable-history-never-mutated invariant (middleware mutates only the per-call ProviderRequest.messages clone; RunOutput.messages returns the untouched durable vector). Enum addition breaks no exhaustive match (checked arcan-lago, arcan-fleet). Tests are substantive, not tautological.

Held below 9 (non-blocking):

  • Feature ships inert — implements arcan_core::runtime::TurnMiddleware (test-only Orchestrator path); production KernelRuntime uses aios-runtime's distinct TurnMiddleware, so a TurnContext adapter is needed to go live. Honestly disclosed in the PR "Scope note" (limitation, not deception).
  • Mild speculative-API slop: compressed_block() + CompressionOutcome public fields have no production consumers yet; compile_context docstring left stale (omits Compressed).

Audit trail per P20; proceeding to p9 auto-merge lifecycle.

@broomva
broomva merged commit 030c416 into main Jul 23, 2026
20 checks passed
@broomva
broomva deleted the feature/bro-425-context-compression-middleware branch July 23, 2026 19:09
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant