Skip to content

Import OpenCode sessions; make the summary model configurable - #1

Closed
artcashin wants to merge 2 commits into
feature/react-frontendfrom
feature/opencode-import-and-model-provider
Closed

artcashin wants to merge 2 commits into
feature/react-frontendfrom
feature/opencode-import-and-model-provider

Conversation

@artcashin

@artcashin artcashin commented Aug 2, 2026 •

Copy link
Copy Markdown
Owner

Adds OpenCode as a session source, and makes the model that writes journal entries configurable. Developed together because the first made the second worth having — a local model is a reasonable choice once the corpus includes local-agent sessions.

Based on feature/react-frontend rather than main so the diff shows only this work; prime-radiant-inc has no branch other than main, so an upstream PR would have carried all 61 React commits. Retarget once prime-radiant-inc#19 lands.

OpenCode source adapter (opt-in)

OpenCode keeps sessions in a single SQLite database rather than one file per session, so it can't be scanned like the Claude Code and Codex sources. The adapter exports each session into a staging directory of JSONL files during ingest, which the normal scanner picks up — dedup, --force, grouping, summarization and the web UI are untouched.

Three things worth a reviewer's attention, each found empirically against OpenCode 1.18.0:

  • opencode session list is scoped to the current directory's project. 5356 bytes from /tmp, 0 bytes from inside another repo. No single cwd can enumerate every project, so sessions are enumerated from OpenCode's database. Transcripts still come from opencode export, which works from anywhere — the fragile part (transcript shape) stays on the supported CLI, and only a five-column query touches the schema.
  • opencode export returns empty stdout over a pipe but is complete when redirected to a file, and one export can run to megabytes. Both call sites go through a single runToFile helper, covered by a 2MB regression test.
  • is_subagent was derived purely from Claude Code's /subagents/ path convention, so flat-staged files always read as 0. Formats that state it outright now win. Claude Code deliberately keeps path detection: its parentSessionId also covers continuations, so switching it to "has a parent" would misflag them.

Projects are keyed off each session's working directory, so OpenCode and Claude Code work in the same repo on the same day lands in one journal entry rather than two disconnected ones.

Configurable summary provider

Summaries and titles were hardcoded to Claude Haiku through the Agent SDK. A summary_provider block now selects any OpenAI-compatible endpoint. Absent config keeps existing behaviour exactly.

  • Both call sites route through the provider. Titles were a second hardcoded Haiku call; leaving them would mean a configured provider silently still hits Anthropic.
  • The Claude auth preflight is skipped for non-Anthropic providers, which authenticate with their own key and would otherwise be blocked on every run.
  • Entries record the model that actually produced them, rather than a constant.
  • API keys are referenced by environment variable name, never stored in config.
  • Reasoning models spend completion tokens before emitting content. An exhausted budget throws with a clear message instead of writing an empty summary.

Testing

221 tests pass, typecheck clean. Every change was written test-first; the OpenAI provider is tested against a real local server rather than a mocked fetch.

Exercised end-to-end against a live corpus: 96 OpenCode sessions across 16 projects (61 subagents correctly flagged), summarized by a local Gemma via llama.cpp. The token-budget error path fired on a real transcript and was fixed by raising the budget — that failure mode is not theoretical.

The web/ React tests are covered separately by vitest, not by bun run check. They were failing under a bare bun test, which is fixed on the base branch (bunfig.toml scopes Bun's runner to src/; CI now runs the web suite with vitest). That fix is merged in here, so a bare bun test on this branch is green.

🤖 Generated with Claude Code

artcashin and others added 2 commits August 2, 2026 13:31
Two additions that were developed together because the first one made the
second one worth having.

OpenCode source adapter
-----------------------
OpenCode keeps sessions in a single SQLite database rather than one file per
session, so it cannot be scanned like the Claude Code and Codex sources. An
opt-in adapter exports each session into a staging directory of JSONL files
during `ingest`, which the normal scanner then picks up — leaving dedup,
--force, grouping, summarization and the web UI untouched.

Three things the implementation has to work around:

- `opencode session list` only reports sessions belonging to the current
  working directory's project, so no single cwd can enumerate them all.
  Sessions are enumerated from OpenCode's database; transcripts still come
  from `opencode export`, which works from anywhere.
- `opencode export` yields empty stdout over a pipe but is complete when
  redirected to a file, and a single export can run to megabytes. Both call
  sites go through one runToFile helper.
- is_subagent was derived purely from Claude Code's /subagents/ path
  convention, so flat-staged files always read as 0. Formats that state it
  outright now win; Claude Code keeps path detection, because its
  parentSessionId also covers continuations and would otherwise misflag them.

Projects are keyed off each session's working directory, so OpenCode and
Claude Code work in the same repo on the same day lands in one journal entry.

Configurable summary provider
-----------------------------
Summaries and session titles were hardcoded to Claude Haiku via the Agent SDK.
A summary_provider config block now selects any OpenAI-compatible endpoint,
including a local llama.cpp server. Absent config keeps the previous behaviour.

Both call sites route through the new provider, since leaving titles on
Anthropic would silently contradict a configured provider, and the Claude auth
preflight is skipped for non-Anthropic providers that authenticate with their
own key. Entries record the model that actually produced them. API keys are
referenced by environment variable name and never stored in config.

Reasoning models spend completion tokens before emitting content, so an
exhausted budget fails loudly rather than writing an empty summary.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@artcashin

Copy link
Copy Markdown
Owner Author

Superseded by prime-radiant-inc#20, which contains the same work branched from main as a standalone PR rather than stacked on feature/react-frontend.

@artcashin artcashin closed this Aug 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant