Repository navigation
Conversation
nnunley
force-pushed
the
feat/recursive-summarization
branch
from
August 13, 2026 01:47
fe9848f to
4aaf3cc
Compare
…rough
Adds a second summary backend alongside the existing Claude path so users
can summarize against any server that speaks the OpenAI API protocol —
Ollama, vLLM, llama.cpp llama-server, LM Studio, OpenAI itself, etc.
Config additions:
- `summary_provider`: `"claude"` (default) or `"openai-compat"`
- `summary_base_url`: e.g. `http://localhost:11434` (path appended)
- `summary_model`: e.g. `qwen3.6:latest`
- `summary_extras`: arbitrary JSON object merged into the request body
The extras passthrough is the key abstraction: rather than enumerating
every server's quirks in code, users put server-specific fields directly
in their config. README documents common values per server (Ollama:
`{"reasoning_effort": "none"}` to suppress reasoning; vLLM/SGLang:
`{"chat_template_kwargs": {"enable_thinking": false}}`; etc.).
Behavioral details:
- `summary_extras` is merged before required fields (model, messages,
stream:false), so users can't accidentally break the request shape.
- Auth preflight (`getClaudeAuthIssue`) is skipped when provider is
`openai-compat`.
- `upsertJournalEntry` now records the actual model used (Claude or
the configured `summary_model`).
- `loadConfig` validates `summary_provider` at load time — invalid
values throw with a clear error listing valid choices.
- New `explainGroup` helper runs the LLM end-to-end without writing
to the DB; intended for audit/debug tools.
Settings UI exposes provider dropdown, base URL, model, and an extras
JSON textarea (with helper text linking common server recipes).
When running `engineering-notebook summarize`, each group's outcome is printed inline as it completes: [1/191] Summarizing moor (2025-09-29)... [1/191] ✓ moor (2025-09-29) [2/191] Summarizing moor (2025-11-14)... [2/191] ⊘ moor (2025-11-14): not engineering work Previously the CLI only emitted the "Summarizing..." line before each group, then dumped all skip reasons in a wall at the end with no way to tell which date+project each reason belonged to. Long --all runs now provide progressive, auditable feedback. Adds an `onComplete` callback to summarizeAll that receives a tagged union outcome (summarized / skipped / error) so callers can render per- group results without re-querying the DB.
…ve fallback
Adds a complete summarization pipeline for the openai-compat provider:
1. One-shot sandwich-format summarization (src/oneshot-summarize.ts).
For transcripts that fit in the model's context budget (default
threshold 800K chars / ~200K tokens for Qwen3.6), wraps the
instructions before AND after the transcript so the format
constraint dominates the model's output style. Uses think:false on
Ollama (10-20x speedup over default thinking) and explicitly lists
required fields in the prompt body since think:false drops schema
requiredness enforcement.
2. Recursive RLM-style orchestrator (src/recursive-summarize.ts) for
transcripts that exceed the one-shot threshold. Strict architectural
invariants:
- Sandboxed scopes: each invocation receives a SCOPE (start, end)
and sees ONLY that slice. Tools (search, read_range, spawn)
operate within scope; children CANNOT expand beyond their
delegated range.
- Spawn shrinkage rules: child range must be a strict subset of
parent's, nonzero, AND at most 50% of parent's length. The 50%
rule prevents trivial-shrink recursion and forces real
divide-and-conquer.
- Spawn forbidden at leaf scopes (≤ READ_WINDOW). Leaves auto-read
their full content. Natural recursion termination at depth
≈ log_2(scope/READ_WINDOW). No artificial invocation cap.
- Mandatory complete coverage: every invocation must cover 100% of
its scope before `done` is honored. Coverage = union of read_range
spans + spawn ranges. Orchestrator computes residual after each
step and shows uncovered ranges in the next plan prompt.
- Coherence-token canary: each invocation generates a per-invocation
random token; model must echo it on `done`. Catches `done` calls
that come from training rather than from this loop's state.
3. Provider-agnostic routing in summarize.ts: per-provider one-shot
threshold (4M chars for Claude, 800K for openai-compat). Above
threshold the recursive orchestrator takes over.
Tests cover: one-shot sandwich invariants, required-field enumeration
in prompts, fitsOneShot threshold logic, scope/span/coverage primitives,
spawn-range validation. The LLM driver itself is exercised by
integration tests in development rather than unit tests because
per-call latency is too high for CI.
Adds a top-level `engineering-notebook skipped` command that lists
journal entries the LLM decided weren't journal-worthy, with the model's
reason text. Useful for spot-checking the skip rate (currently ~50% of
day-project groups in our corpus).
Filters: --since YYYY-MM-DD, --project NAME, --limit N (default 20).
Each entry shows: date, project, skip reason text, and generated_at
timestamp. To re-evaluate an entry the user thinks was wrongly skipped,
the help text describes the manual SQL DELETE + summarize re-run flow.
A future change could add a Web UI surface (a /skipped page in the
notebook server). For now the CLI is the audit interface.
Hypothesis-forming on current skip data (2026-05-13, 8 recent skips):
- All 8 had tiny day-scoped content (<2KB; smallest 389 chars).
- The 398-char zerocopy/boxter dups are literal `/exit` boilerplate
where the user just typed /exit at end of day. Skips are correct.
- Larger ones (mmmmyes 1075c, 9front 1484c) might be legitimate
false-negatives where short planning/inspection arcs got rejected
for "no engineering work was completed or shipped" — the model is
biased toward "shipping" as the engineering bar.
Adds a robust LLM JSON response parser that handles common failure modes
observed during the recursive orchestrator's smoke tests:
1. Markdown code fences (\`\`\`json ... \`\`\`) — already handled, kept.
2. <think>...</think> leaks: some providers (notably Ollama with
thinking-capable models) emit reasoning tokens inline in content even
when think:false is set. Strip both balanced <think>...</think> blocks
and dangling </think> tags.
3. Prose preamble: model says "Sure! Here is your JSON: {...}". Recover
by finding the first balanced {...} substring (respecting string
literals and escape sequences).
The new `parseLlmJsonResponse` and `extractFirstBalancedObject` helpers
are exported and used by both recursive-summarize.ts and
oneshot-summarize.ts for consistent recovery semantics across all
provider paths.
Reduces forced-extract rate during recursive runs (observed in spike
testing on 492KB transcript: 3 of 13 invocations had non-JSON failures
under the old parser; the new parser recovers from at least the
<think> leak case).
Tests cover: clean JSON, fenced JSON, <think> leaks, prose preamble,
unterminated input, top-level non-objects, balanced-brace scanning
with strings and escapes.
Previously, sessions whose messages spanned midnight were split across
two date buckets using per-message logical dates. In practice this
created thin tail slivers on the next day — a few morning messages from
a session that started the prior evening — that the LLM consistently
rejected as "no engineering work" because the slivers lacked context.
Skip-pattern analysis (2026-05-13, 8 recent skips) found that all
skipped groups had tiny day-scoped content (<2KB; smallest 389 chars
literal `/exit` boilerplate). The most recoverable false-negatives
were these midnight-split slivers where substantive work happened on
the start side of midnight but the tail got rejected in isolation.
Fix: each session is now atomic and attributed to its `started_at`
logical date (the date adjusted by `day_start_hour`). A session that
begins late one night and continues past midnight contributes its full
content to the start date.
`splitConversationByDay` is kept as an exported helper for external
tools that want per-message day attribution, but the orchestrator no
longer uses it.
Tests updated: the midnight-spanning suite now verifies atomic
attribution (one group, all 6 messages on Feb 20) and that
`filterDate("2026-02-21")` returns nothing for a session that started
on Feb 20 even when some messages are timestamped Feb 21.
Adds a journal_invocations table that mirrors the recursion tree of an RLM-style summarization run, plus an FK from journal_fragments to its producing invocation. The recursive orchestrator now records every action it takes (search, read_range, spawn, done, done-rejected, plan-failed, extract-failed) and each invocation's full action timeline is persisted as a JSON column. Fragments carry the source span and SHA-1 checksum of their content at extraction time. Schema: - journal_invocations (id, parent_id, journal_entry_id, scope_start, scope_end, depth, coherence_token, question, actions JSON, started_at, ended_at). parent_id self-references for the recursion tree; journal_entry_id has ON DELETE CASCADE. - journal_fragments gains invocation_id INTEGER REFERENCES journal_invocations(id) ON DELETE CASCADE. The migration uses the table-rebuild dance (CREATE+COPY+DROP+RENAME) since SQLite forbids ALTER TABLE ADD COLUMN with REFERENCES. API additions: - summarizeRecursive now returns rootInvocation: InvocationNode alongside the parsed result. - summarizeRecursive accepts options.sessionRanges for source attribution of fragments back to specific session_ids. - persistInvocationTree(db, journal_entry_id, root) writes the tree. - loadInvocationTree(db, journal_entry_id) reconstructs it. - renderInvocationTree(tree) produces a human-readable trace for --why. CLI changes: - summarize --why first looks for a stored invocation tree and renders it. Falls back to the previous live-re-run via explainGroup if no tree exists (which is the case for one-shot entries; only recursive runs produce trees). The recursion tree gives --why the full audit story: every action the model took, every span it inspected, every child it spawned, every extraction it produced. The fragments table generalizes from the old chunked map-reduce role into a per-invocation evidence ledger. Tests cover: round-trip persistence (single, nested, idempotent rewrite), action recording, render output. 16 new tests; 168 total.
nnunley
force-pushed
the
feat/recursive-summarization
branch
from
August 13, 2026 01:51
4aaf3cc to
f6f9eef
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Draft / proposal. This is offered, not assumed. It is a subsystem, not a patch: new config keys, three new tables, two new CLI verbs, and a ~1,200-line orchestrator. If the maintainers do not want to own this surface, say so and it stays downstream. If parts of it are wanted and parts are not, it can be cut down.
Problem
Some days produce transcripts larger than any single context window. The existing path builds one prompt per (date, project) group and sends it to the model. When the group is too large, the options today are truncation or a skip — both silently lose engineering history, and a skip is indistinguishable from a day with no work.
A naive map-reduce over chunks does not solve it either, because nothing guarantees the model actually looked at every chunk. A summarizer that quietly covers 60% of a day and returns a confident summary is worse than one that fails loudly.
Approach
summarizepicks a strategy by transcript size:src/oneshot-summarize.ts) when the transcript fits under the provider threshold. Sandwich prompt with the instructions repeated before and after the transcript, plus explicit anti-continuation guidance so the model summarizes instead of extending the transcript.src/recursive-summarize.ts) otherwise.The recursive summarizer gives the model a scope — a byte range of the transcript — and three tools:
search,read_range, andspawn.spawncreates a child invocation over a sub-range, which recurses with the same rules. Each invocation returns a structured extraction that is folded into its parent.The 100% coverage guarantee
An invocation may only declare
donewhen the union of itsread_rangespans, plus the scopes of its children, covers its entire scope. Coverage is tracked as span algebra:mergeSpansnormalizes and unions overlapping and adjacent spans.residualGapscomputes the uncovered intervals of[0, scopeLen).donewith non-empty residual gaps is rejected, and the gaps are handed back to the model as the remaining work.This is why it matters: the failure mode of every LLM summarizer over long inputs is confident partial coverage. Here, partial coverage cannot be reported as success. The summarizer can fail — it cannot silently under-read.
Termination comes from
validateSpawnRange: a child scope must be in bounds, non-empty, and at most 50% of its parent's length. Scope therefore at least halves per level, so recursion depth is bounded bylog2of transcript size. Each invocation is additionally capped atMAX_ACTIONS_PER_INVOCATION(20) andread_rangereturns at mostREAD_WINDOW(2000) bytes per call, so a single invocation cannot brute-force a large scope without spawning.Local models return dirty JSON.
parseLlmJsonResponsestrips ```json and bare fences, strips<think>...</think>leaks including a dangling closing tag with no opener, and falls back to `extractFirstBalancedObject`, a brace matcher that respects string literals and escapes. Top-level arrays and scalars are rejected rather than coerced.Auditability
The invocation tree is persisted (
journal_invocations,journal_fragments) with per-invocation scope, actions, and outcome, and rendered byrenderInvocationTree.summarize --whyexplains why a group produced the entry it did, and a newskippedverb audits skip reasons across the corpus. Without this the recursion is a black box; with it, an under-covered or misbehaving run is inspectable after the fact.Supporting pieces in this branch
summary_provider,summary_base_url,summary_model,summary_extras), with extras passthrough for provider-specific knobs. Required request keys override extras so extras cannot break the request shape. Config validation rejects an unknown provider with a clear error. Settings are exposed in the web UI.src/rollup.ts,journal_rollups,rollupverb) aggregating daily entries into a Monday-anchored week, keeping skipped days visible in a separate section.Scope note
This branch also contains the two ingest/grouping fixes submitted separately as #27. If #27 merges first, this branch rebases onto it and those commits drop out. Reviewing #27 first is the easier path.
How to test
166 passing, 0 failing, typecheck clean (baseline on
mainis 92 passing).Unit coverage is on the parts where correctness is checkable without a model:
mergeSpans,residualGaps,validateSpawnRange(including the 51%-of-parent rejection),extractFirstBalancedObject,parseLlmJsonResponse(including a real Ollama think-leak case), and invocation-tree round-trip, idempotent rewrite, and rendering.Known limitations
IF NOT EXISTSand guardedALTER TABLE, butjournal_fragmentsis rebuilt via table-rename to add a foreign key, since SQLite cannot add a column withREFERENCES. That path deserves attention in review.