Skip to content

Proposal: recursive summarization for transcripts larger than the context window - #28

Draft
nnunley wants to merge 10 commits into
prime-radiant-inc:mainfrom
nnunley:feat/recursive-summarization
Draft

nnunley wants to merge 10 commits into
prime-radiant-inc:mainfrom
nnunley:feat/recursive-summarization

Conversation

@nnunley

@nnunley nnunley commented Aug 13, 2026

Copy link
Copy Markdown

Draft / proposal. This is offered, not assumed. It is a subsystem, not a patch: new config keys, three new tables, two new CLI verbs, and a ~1,200-line orchestrator. If the maintainers do not want to own this surface, say so and it stays downstream. If parts of it are wanted and parts are not, it can be cut down.

Problem

Some days produce transcripts larger than any single context window. The existing path builds one prompt per (date, project) group and sends it to the model. When the group is too large, the options today are truncation or a skip — both silently lose engineering history, and a skip is indistinguishable from a day with no work.

A naive map-reduce over chunks does not solve it either, because nothing guarantees the model actually looked at every chunk. A summarizer that quietly covers 60% of a day and returns a confident summary is worse than one that fails loudly.

Approach

summarize picks a strategy by transcript size:

  • One-shot fast path (src/oneshot-summarize.ts) when the transcript fits under the provider threshold. Sandwich prompt with the instructions repeated before and after the transcript, plus explicit anti-continuation guidance so the model summarizes instead of extending the transcript.
  • Recursive fallback (src/recursive-summarize.ts) otherwise.

The recursive summarizer gives the model a scope — a byte range of the transcript — and three tools: search, read_range, and spawn. spawn creates a child invocation over a sub-range, which recurses with the same rules. Each invocation returns a structured extraction that is folded into its parent.

The 100% coverage guarantee

An invocation may only declare done when the union of its read_range spans, plus the scopes of its children, covers its entire scope. Coverage is tracked as span algebra:

  • mergeSpans normalizes and unions overlapping and adjacent spans.
  • residualGaps computes the uncovered intervals of [0, scopeLen).
  • A done with non-empty residual gaps is rejected, and the gaps are handed back to the model as the remaining work.

This is why it matters: the failure mode of every LLM summarizer over long inputs is confident partial coverage. Here, partial coverage cannot be reported as success. The summarizer can fail — it cannot silently under-read.

Termination comes from validateSpawnRange: a child scope must be in bounds, non-empty, and at most 50% of its parent's length. Scope therefore at least halves per level, so recursion depth is bounded by log2 of transcript size. Each invocation is additionally capped at MAX_ACTIONS_PER_INVOCATION (20) and read_range returns at most READ_WINDOW (2000) bytes per call, so a single invocation cannot brute-force a large scope without spawning.

Local models return dirty JSON. parseLlmJsonResponse strips ```json and bare fences, strips <think>...</think> leaks including a dangling closing tag with no opener, and falls back to `extractFirstBalancedObject`, a brace matcher that respects string literals and escapes. Top-level arrays and scalars are rejected rather than coerced.

Auditability

The invocation tree is persisted (journal_invocations, journal_fragments) with per-invocation scope, actions, and outcome, and rendered by renderInvocationTree. summarize --why explains why a group produced the entry it did, and a new skipped verb audits skip reasons across the corpus. Without this the recursion is a black box; with it, an under-covered or misbehaving run is inspectable after the fact.

Supporting pieces in this branch

  • OpenAI-compatible provider alongside the existing Claude path (summary_provider, summary_base_url, summary_model, summary_extras), with extras passthrough for provider-specific knobs. Required request keys override extras so extras cannot break the request shape. Config validation rejects an unknown provider with a clear error. Settings are exposed in the web UI.
  • Weekly per-project rollups (src/rollup.ts, journal_rollups, rollup verb) aggregating daily entries into a Monday-anchored week, keeping skipped days visible in a separate section.

Scope note

This branch also contains the two ingest/grouping fixes submitted separately as #27. If #27 merges first, this branch rebases onto it and those commits drop out. Reviewing #27 first is the easier path.

How to test

bun test
bun run typecheck

166 passing, 0 failing, typecheck clean (baseline on main is 92 passing).

Unit coverage is on the parts where correctness is checkable without a model: mergeSpans, residualGaps, validateSpawnRange (including the 51%-of-parent rejection), extractFirstBalancedObject, parseLlmJsonResponse (including a real Ollama think-leak case), and invocation-tree round-trip, idempotent rewrite, and rendering.

Known limitations

  • The 100% guarantee is over byte coverage, not comprehension. It proves every region was read by some invocation; it does not prove the extraction from that region was good.
  • Recursion costs many model calls on large transcripts. It is a fallback, not the normal path — the one-shot path handles typical days.
  • Tuning constants (50% spawn cap, 20 actions, 2000-byte read window, provider size thresholds) are hardcoded, chosen empirically against local models, and not yet configurable.
  • End-to-end recursion is not covered by automated tests because it requires a live model; the deterministic components are unit-tested and the orchestration loop is not.
  • New config keys default to the existing Claude behavior, so an existing install is unaffected until it opts in.
  • Schema migrations are additive with IF NOT EXISTS and guarded ALTER TABLE, but journal_fragments is rebuilt via table-rename to add a foreign key, since SQLite cannot add a column with REFERENCES. That path deserves attention in review.

@nnunley
nnunley force-pushed the feat/recursive-summarization branch from fe9848f to 4aaf3cc Compare August 13, 2026 01:47
nnunley added 10 commits August 12, 2026 21:51
…rough

Adds a second summary backend alongside the existing Claude path so users
can summarize against any server that speaks the OpenAI API protocol —
Ollama, vLLM, llama.cpp llama-server, LM Studio, OpenAI itself, etc.

Config additions:
- `summary_provider`: `"claude"` (default) or `"openai-compat"`
- `summary_base_url`: e.g. `http://localhost:11434` (path appended)
- `summary_model`: e.g. `qwen3.6:latest`
- `summary_extras`: arbitrary JSON object merged into the request body

The extras passthrough is the key abstraction: rather than enumerating
every server's quirks in code, users put server-specific fields directly
in their config. README documents common values per server (Ollama:
`{"reasoning_effort": "none"}` to suppress reasoning; vLLM/SGLang:
`{"chat_template_kwargs": {"enable_thinking": false}}`; etc.).

Behavioral details:
- `summary_extras` is merged before required fields (model, messages,
  stream:false), so users can't accidentally break the request shape.
- Auth preflight (`getClaudeAuthIssue`) is skipped when provider is
  `openai-compat`.
- `upsertJournalEntry` now records the actual model used (Claude or
  the configured `summary_model`).
- `loadConfig` validates `summary_provider` at load time — invalid
  values throw with a clear error listing valid choices.
- New `explainGroup` helper runs the LLM end-to-end without writing
  to the DB; intended for audit/debug tools.

Settings UI exposes provider dropdown, base URL, model, and an extras
JSON textarea (with helper text linking common server recipes).
When running `engineering-notebook summarize`, each group's outcome is
printed inline as it completes:

  [1/191] Summarizing moor (2025-09-29)...
  [1/191] ✓ moor (2025-09-29)
  [2/191] Summarizing moor (2025-11-14)...
  [2/191] ⊘ moor (2025-11-14): not engineering work

Previously the CLI only emitted the "Summarizing..." line before each
group, then dumped all skip reasons in a wall at the end with no way to
tell which date+project each reason belonged to. Long --all runs now
provide progressive, auditable feedback.

Adds an `onComplete` callback to summarizeAll that receives a tagged
union outcome (summarized / skipped / error) so callers can render per-
group results without re-querying the DB.
…ve fallback

Adds a complete summarization pipeline for the openai-compat provider:

1. One-shot sandwich-format summarization (src/oneshot-summarize.ts).
   For transcripts that fit in the model's context budget (default
   threshold 800K chars / ~200K tokens for Qwen3.6), wraps the
   instructions before AND after the transcript so the format
   constraint dominates the model's output style. Uses think:false on
   Ollama (10-20x speedup over default thinking) and explicitly lists
   required fields in the prompt body since think:false drops schema
   requiredness enforcement.

2. Recursive RLM-style orchestrator (src/recursive-summarize.ts) for
   transcripts that exceed the one-shot threshold. Strict architectural
   invariants:

   - Sandboxed scopes: each invocation receives a SCOPE (start, end)
     and sees ONLY that slice. Tools (search, read_range, spawn)
     operate within scope; children CANNOT expand beyond their
     delegated range.

   - Spawn shrinkage rules: child range must be a strict subset of
     parent's, nonzero, AND at most 50% of parent's length. The 50%
     rule prevents trivial-shrink recursion and forces real
     divide-and-conquer.

   - Spawn forbidden at leaf scopes (≤ READ_WINDOW). Leaves auto-read
     their full content. Natural recursion termination at depth
     ≈ log_2(scope/READ_WINDOW). No artificial invocation cap.

   - Mandatory complete coverage: every invocation must cover 100% of
     its scope before `done` is honored. Coverage = union of read_range
     spans + spawn ranges. Orchestrator computes residual after each
     step and shows uncovered ranges in the next plan prompt.

   - Coherence-token canary: each invocation generates a per-invocation
     random token; model must echo it on `done`. Catches `done` calls
     that come from training rather than from this loop's state.

3. Provider-agnostic routing in summarize.ts: per-provider one-shot
   threshold (4M chars for Claude, 800K for openai-compat). Above
   threshold the recursive orchestrator takes over.

Tests cover: one-shot sandwich invariants, required-field enumeration
in prompts, fitsOneShot threshold logic, scope/span/coverage primitives,
spawn-range validation. The LLM driver itself is exercised by
integration tests in development rather than unit tests because
per-call latency is too high for CI.
Adds a top-level `engineering-notebook skipped` command that lists
journal entries the LLM decided weren't journal-worthy, with the model's
reason text. Useful for spot-checking the skip rate (currently ~50% of
day-project groups in our corpus).

Filters: --since YYYY-MM-DD, --project NAME, --limit N (default 20).

Each entry shows: date, project, skip reason text, and generated_at
timestamp. To re-evaluate an entry the user thinks was wrongly skipped,
the help text describes the manual SQL DELETE + summarize re-run flow.

A future change could add a Web UI surface (a /skipped page in the
notebook server). For now the CLI is the audit interface.

Hypothesis-forming on current skip data (2026-05-13, 8 recent skips):
  - All 8 had tiny day-scoped content (<2KB; smallest 389 chars).
  - The 398-char zerocopy/boxter dups are literal `/exit` boilerplate
    where the user just typed /exit at end of day. Skips are correct.
  - Larger ones (mmmmyes 1075c, 9front 1484c) might be legitimate
    false-negatives where short planning/inspection arcs got rejected
    for "no engineering work was completed or shipped" — the model is
    biased toward "shipping" as the engineering bar.
Adds a robust LLM JSON response parser that handles common failure modes
observed during the recursive orchestrator's smoke tests:

1. Markdown code fences (\`\`\`json ... \`\`\`) — already handled, kept.

2. <think>...</think> leaks: some providers (notably Ollama with
   thinking-capable models) emit reasoning tokens inline in content even
   when think:false is set. Strip both balanced <think>...</think> blocks
   and dangling </think> tags.

3. Prose preamble: model says "Sure! Here is your JSON: {...}". Recover
   by finding the first balanced {...} substring (respecting string
   literals and escape sequences).

The new `parseLlmJsonResponse` and `extractFirstBalancedObject` helpers
are exported and used by both recursive-summarize.ts and
oneshot-summarize.ts for consistent recovery semantics across all
provider paths.

Reduces forced-extract rate during recursive runs (observed in spike
testing on 492KB transcript: 3 of 13 invocations had non-JSON failures
under the old parser; the new parser recovers from at least the
<think> leak case).

Tests cover: clean JSON, fenced JSON, <think> leaks, prose preamble,
unterminated input, top-level non-objects, balanced-brace scanning
with strings and escapes.
Previously, sessions whose messages spanned midnight were split across
two date buckets using per-message logical dates. In practice this
created thin tail slivers on the next day — a few morning messages from
a session that started the prior evening — that the LLM consistently
rejected as "no engineering work" because the slivers lacked context.

Skip-pattern analysis (2026-05-13, 8 recent skips) found that all
skipped groups had tiny day-scoped content (<2KB; smallest 389 chars
literal `/exit` boilerplate). The most recoverable false-negatives
were these midnight-split slivers where substantive work happened on
the start side of midnight but the tail got rejected in isolation.

Fix: each session is now atomic and attributed to its `started_at`
logical date (the date adjusted by `day_start_hour`). A session that
begins late one night and continues past midnight contributes its full
content to the start date.

`splitConversationByDay` is kept as an exported helper for external
tools that want per-message day attribution, but the orchestrator no
longer uses it.

Tests updated: the midnight-spanning suite now verifies atomic
attribution (one group, all 6 messages on Feb 20) and that
`filterDate("2026-02-21")` returns nothing for a session that started
on Feb 20 even when some messages are timestamped Feb 21.
Adds a journal_invocations table that mirrors the recursion tree of an
RLM-style summarization run, plus an FK from journal_fragments to its
producing invocation. The recursive orchestrator now records every
action it takes (search, read_range, spawn, done, done-rejected,
plan-failed, extract-failed) and each invocation's full action timeline
is persisted as a JSON column. Fragments carry the source span and
SHA-1 checksum of their content at extraction time.

Schema:
- journal_invocations (id, parent_id, journal_entry_id, scope_start,
  scope_end, depth, coherence_token, question, actions JSON, started_at,
  ended_at). parent_id self-references for the recursion tree;
  journal_entry_id has ON DELETE CASCADE.
- journal_fragments gains invocation_id INTEGER REFERENCES
  journal_invocations(id) ON DELETE CASCADE. The migration uses the
  table-rebuild dance (CREATE+COPY+DROP+RENAME) since SQLite forbids
  ALTER TABLE ADD COLUMN with REFERENCES.

API additions:
- summarizeRecursive now returns rootInvocation: InvocationNode
  alongside the parsed result.
- summarizeRecursive accepts options.sessionRanges for source
  attribution of fragments back to specific session_ids.
- persistInvocationTree(db, journal_entry_id, root) writes the tree.
- loadInvocationTree(db, journal_entry_id) reconstructs it.
- renderInvocationTree(tree) produces a human-readable trace for --why.

CLI changes:
- summarize --why first looks for a stored invocation tree and renders
  it. Falls back to the previous live-re-run via explainGroup if no
  tree exists (which is the case for one-shot entries; only recursive
  runs produce trees).

The recursion tree gives --why the full audit story: every action the
model took, every span it inspected, every child it spawned, every
extraction it produced. The fragments table generalizes from the old
chunked map-reduce role into a per-invocation evidence ledger.

Tests cover: round-trip persistence (single, nested, idempotent
rewrite), action recording, render output. 16 new tests; 168 total.
@nnunley
nnunley force-pushed the feat/recursive-summarization branch from 4aaf3cc to f6f9eef Compare August 13, 2026 01:51
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant