Skip to content

Weekly Status Reports: aggregate a week's entries into a manager-facing report - #25

Draft
artcashin wants to merge 81 commits into
prime-radiant-inc:mainfrom
artcashin:feat/weekly-status-reports
Draft

artcashin wants to merge 81 commits into
prime-radiant-inc:mainfrom
artcashin:feat/weekly-status-reports

Conversation

@artcashin

Copy link
Copy Markdown

Aggregates a week's journal entries into a status report a developer can send to their manager — narrative, accomplishments, per-project breakdown, and items still outstanding. Generated from a user-editable prompt template, stored with version history, exported as markdown. Triggered from the CLI, a weekly launchd job, or a button in the web UI.

Draft — stacked. This branch is based on the OpenCode/model-provider work, so its diff against main currently shows ~79 commits. Only 13 are this feature. It contains #19 and #22; it needs llm.ts from #20. The diff collapses to those 13 commits once those merge. Review commits from 75d0b0c onward, or wait for the base to land.

Why the prompt is a file, not a constant

Different managers want different things, and the person best placed to say what a status report should contain is the manager receiving it. report_template_url points at a markdown file; edit it once and a whole team's reports change shape. Placeholders: {{week_label}}, {{week_start}}, {{week_end}}, {{entries}}, {{open_questions}}, {{project_list}}. Unknown placeholders are left untouched, so a typo degrades one line rather than killing the run.

Every successful fetch is cached. If the URL is unreachable the cached copy is used. If there is no cache, generation fails rather than silently falling back to the shipped default — a differently-shaped report going out under someone's name is worse than no report.

Design decisions worth a reviewer's attention

  • Versioned, never overwritten. A report may already have been sent, so the record of what was sent matters. Each generation is a new row; latest wins by version DESC.
  • Markdown is the source of truth. A ## heading parse sits on top as a convenience and is allowed to come back empty — a user-authored template is under no obligation to use headings.
  • Outstanding items are judged, not just collected. The model weighs open questions against the week's later entries and drops ones that were resolved, then lists what it dropped under "Resolved This Week" so the judgement is auditable.
  • Single model call. The busiest measured week is ~11k characters; chunking would be premature.
  • 05:15 Monday, not midnight. With day_start_hour at 5 the previous logical week does not close until 05:00 Monday, so a midnight job would report on an incomplete week and freeze it that way.

Verified against real data, not just tests

265 tests pass, typecheck clean. Beyond that:

  • A live model call produced a real five-section report for 2026-W31 from actual journal entries.
  • Both trigger paths exercised independently — v1 from the CLI, v2 from the web button — which also demonstrates versioning appends rather than overwrites.
  • The launchd job is installed and was triggered: it computed the last completed week on its own, ran under launchd's minimal environment, exited 0, wrote nothing to stderr.
  • The template URL chain was exercised against a real HTTP server in all four states: fetch succeeds and caches; server down with cache falls back; server down without cache fails hard with exit 1; and a custom remote template visibly reshaped the model's output. The database records which was used per report (template_source).

Known limitations, stated rather than discovered

  • resolveTemplate has no fetch timeout. A URL that accepts the connection and never responds blocks the run. AbortSignal.timeout() would close it.
  • The status endpoint returns HTML, so the React button in the follow-up plan will need JSON or a second route.
  • Journal text is interpolated into the prompt without delimiting. Those strings are LLM output over session transcripts that can contain third-party text; this is the first feature that forwards them to someone else.
  • The job registry is an in-memory Map, deliberately. The report commits before the job is marked done, so a restart loses job status but never a report.

The React button is not in this PR. It depends on the React frontend and is a separate plan.

🤖 Generated with Claude Code

artcashin and others added 30 commits July 19, 2026 07:58
Manual groups over coding sessions: a Groups nav tab, one group per
session, in-app create/rename/delete, membership in a separate table
that survives re-ingest.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
7 TDD tasks: schema, helper module, nav, views, routes, session
control, verification. Corrects spec FK note (foreign_keys=ON means
session_id must have no FK to survive re-ingest).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
One-way import of Claude Desktop groups (dframe-group-scopes) into the
notebook: reader + reconcile (mirror by desktop_id, manual groups
untouched), Ungrouped view, and a block-while-Desktop-open guard behind
a single policy switch. Write-back deferred to Phase 2.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
7 TDD tasks: desktop_id migration, classic-level reader (schema pinned
against real store), reconcile/mirror, Ungrouped view, block-while-open
guard, import route+button+CLI, verification.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Approach A (re-parse source JSONL on demand). Default text-only; Show
thinking/tools via query params re-parses source_path, warns if the
file is unavailable. No ingest/schema change. Precedes the group work.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
4 TDD tasks: transcript parser, enriched session render + toggle
controls + CSS, route params, verification. Re-parses source_path on
demand; default view unchanged.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…es, param-preservation test

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Core match: collapsible tool calls (name + preview, result paired
inside), thinking bubbles + token estimate, teal accent + rounded
styling on the notebook's light palette. No markdown/diff (deferred),
no new deps.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
3 TDD tasks: parser tool id/input enrichment, collapsible-tool +
thinking-bubble render rework + CSS, verification.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…t edge

Final-review cleanups: remove redundant style="display:block" on the
orphan tool_result div; update spec edge-case wording to match the
implemented (extras-dropped) behavior.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Umbrella doc: keep engineering-notebook backend, new React/Vite frontend
matching claude-session-viewer's session display, plus Desktop groups.
Recommends monorepo + a 5-phase roadmap (Phase 1 = session display
parity). Evidence-based session-chain findings (resume = same session;
subagents are the multi-file case). Each phase gets its own spec/plan.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
React/Vite app in web/, Hono JSON API (session list, structured
transcript, subagent), polished session viewer (collapsible tools,
thinking, subagent nesting, client-side hide/show toggles), Tailwind.
Subagent->Task mapping via first-prompt match (viewer's strategy).
Spikes deferred to plan: on-disk subagent layout, Vite<->Hono wiring.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
7 tasks: Vite/Hono scaffold+wiring, structured transcript endpoint,
subagent discovery via .meta.json + endpoints, session list, React API
layer + list view, session viewer (collapsible tools/thinking/subagents
+ toggles), verification. Additive-only; legacy views/DB untouched.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- web/ Vite React-TS + Tailwind SPA; App pings /api
- createApiRouter mounted at /api; GET /api/ping
- createApp({react}) serves web/dist SPA (assets + index.html fallback),
  skipping legacy views; --react flag threaded through serve
- exclude web/ from backend tsconfig
Verified end-to-end: dev proxy + prod build; /api/ping, SPA, assets, and
client-route fallback all 200.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
artcashin and others added 30 commits July 19, 2026 23:36
Ingested subagent transcripts (<project>/<parentUUID>/subagents/agent-*.jsonl)
now record parent_session_id — taken from the records' own sessionId without
triggering the continuation prefix-skip (a separate continuationParentId drives
that). A new relinkSubagents() fills the link (from the column or recovered from
the path) and re-homes each subagent's project_id to its parent's project. It
runs at the end of ingest and is order-independent + idempotent, so it also
backfills. Adds idx_sessions_parent.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
/api/sessions/:id now returns the subagent backlink (parent_title, gated on
is_subagent so continuations aren't mislabeled), subtask_title, spawn_tool_use_id,
and the session's ingested subagents ordered chronologically. Journal, project,
and group rows carry their nested subagents via a shared ingestedSubagents()
helper (single batched query, filtered to ingested rows so there are no dead
links). Ungrouped and the flat /api/sessions list exclude subagents.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Panel 2 nests each session's subagents (revealed once the session is selected),
labelled by subtask title. Panel 3 gains: a compaction-proof SubagentIndex; a
"Subagent of …" backlink (shown only for subagents) that navigates to the parent
and scrolls/ring-highlights the exact spawn point, falling back to the index when
the spawn was compacted off the transcript; inline spawn links matched by
tool_use id (name-agnostic — Claude Code names the tool "Agent", not "Task"),
visible even with tools hidden; and subtask titles as the panel header. The
focus scroll fires once per target and the window grows-only, so reading past a
spawn no longer snaps back.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The date was interpolated raw into the muted/error summary cards; wrap it in
escapeHtml() like the surrounding fields.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Ingestion was nested inside the remoteSources block, so a purely local setup
never scanned/ingested on sync. Hoist the scanSources/ingestSessions call so it
always runs, appending any remote cached paths when present.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
… toggle

Compacted sessions keep their full pre-compaction messages in the same file, but
active-path resolution prunes them (e.g. 1,036 shown of 3,122). Panel 3 now
defaults to the uncompacted view: the transcript parser gains a `full` option
(returns every message in file order, dropping the injected isCompactSummary
blob) and a `compacted` flag; the endpoint honors ?full=1. A third top-bar
toggle — "View compacted" / "View uncompacted" — switches modes, defaulting to
uncompacted so pre-compaction detail (and its subagent spawn links) is shown.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
… compacted

Plumb the transcript's `compacted` flag to the toggle bar. The "View compacted"
button is disabled (and greyed) unless the open session actually has a compacted
view to switch to. Non-compacted sessions are forced back to the default
uncompacted view so the disabled toggle never strands the reader in compacted
mode; the toggle re-disables when leaving the session.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Light green (matching the other toggles) when the session has a compacted view;
gray when it doesn't.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…uped to remove)

Adds POST /api/sessions/:id/group ({ groupId } assigns, { groupId: null }
unassigns; 404s on missing session/group). In the Groups view, panel-2 session
rows are draggable and panel-1 group rows (and Ungrouped) are drop targets with a
hover highlight; dropping assigns/unassigns the conversation and refreshes the
counts and current list. Subagent rows stay non-draggable.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Two additions that were developed together because the first one made the
second one worth having.

OpenCode source adapter
-----------------------
OpenCode keeps sessions in a single SQLite database rather than one file per
session, so it cannot be scanned like the Claude Code and Codex sources. An
opt-in adapter exports each session into a staging directory of JSONL files
during `ingest`, which the normal scanner then picks up — leaving dedup,
--force, grouping, summarization and the web UI untouched.

Three things the implementation has to work around:

- `opencode session list` only reports sessions belonging to the current
  working directory's project, so no single cwd can enumerate them all.
  Sessions are enumerated from OpenCode's database; transcripts still come
  from `opencode export`, which works from anywhere.
- `opencode export` yields empty stdout over a pipe but is complete when
  redirected to a file, and a single export can run to megabytes. Both call
  sites go through one runToFile helper.
- is_subagent was derived purely from Claude Code's /subagents/ path
  convention, so flat-staged files always read as 0. Formats that state it
  outright now win; Claude Code keeps path detection, because its
  parentSessionId also covers continuations and would otherwise misflag them.

Projects are keyed off each session's working directory, so OpenCode and
Claude Code work in the same repo on the same day lands in one journal entry.

Configurable summary provider
-----------------------------
Summaries and session titles were hardcoded to Claude Haiku via the Agent SDK.
A summary_provider config block now selects any OpenAI-compatible endpoint,
including a local llama.cpp server. Absent config keeps the previous behaviour.

Both call sites route through the new provider, since leaving titles on
Anthropic would silently contradict a configured provider, and the Claude auth
preflight is skipped for non-Anthropic providers that authenticate with their
own key. Entries record the model that actually produced them. API keys are
referenced by environment variable name and never stored in config.

Reasoning models spend completion tokens before emitting content, so an
exhausted budget fails loudly rather than writing an empty summary.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`bun test` collected web/src/**/*.test.tsx and failed 18 of them with
"document is not defined". Those are React component tests that need a DOM,
which is configured for vitest in web/vite.config.ts — a file Bun's test
runner does not read. Nothing was wrong with the tests themselves: they pass
under their intended runner.

The failure was invisible locally because the pre-commit hook runs the scoped
`bun test ./src`, while both CI workflows run a bare `bun test`. main has no
web tests, so this branch would have turned CI red on merge.

bunfig.toml scopes Bun's runner to src/, leaving it responsible for the
backend suite alone. That also excludes web/src/session/adapt.test.ts, which
is DOM-free and did run under Bun, so CI now runs the web suite with vitest —
covering all 20 web tests, which CI never executed at all. web/ is not a
workspace of the root package, so its dependencies install separately.

Verified: bare `bun test` 182 pass / 0 fail, `bun run test:web` 20 pass,
typecheck clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A day is summarized long before it is over. Because group selection filtered
out every (date, project) that already had an entry — and the --date filter was
applied inside that guard — no flag could refresh the current day's entry.
Neither --all nor --date reached it; the only way was deleting the row by hand.

Naming a date is now treated as a deliberate request to regenerate it. --all is
unchanged and still only visits dates with no entry, so routine incremental runs
never rewrite existing work. Persistence was already an upsert, so a regenerated
entry replaces the old one in place.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Brings the explicit-date regeneration into the branch the local install runs
from, so a scheduled nightly job can refresh the previous day's entries.

Conflict was in summarize.test.ts, where both branches appended a describe
block; resolution keeps both.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Aggregates a week's journal entries into a manager-facing report: narrative,
accomplishments, per-project breakdown, and outstanding items judged against
the week's later entries so mid-week resolutions are not reported as open.

The prompt is a fully user-editable markdown template fetched from a URL, so a
manager can set the format for a whole team without anyone editing TypeScript.
Markdown is the stored source of truth; a heading parse sits on top as a
convenience that cannot break generation.

Every generation is a new version rather than an overwrite, because a report
may already have been sent and the record of what was sent matters.

Design settled through brainstorming; no open questions.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Seven TDD tasks covering week math, template resolution, schema and config,
the generation core, the CLI with markdown export, the HTMX view with async
job polling, and the scheduled launchd job.

Week-boundary test expectations are computed from real ISO calendar values,
including the two cases most likely to be got wrong: 2026-01-01 belongs to
W01 whose week begins 2025-12-29, and 2026 has a W53 spilling into 2027.

Task 1 supersedes the weekMonday helper currently in calendar.ts so week math
lives in one place rather than two.

The React button is deferred to its own plan; it depends on the unmerged
React frontend and cannot ship to main independently.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Pre-flight review found three places the plan mandated code the review rubric
treats as a defect: an `as any` return, a bare catch that discarded the
original error before retrying, and duplicated prompt assembly between the CLI's
--stdout and normal paths. Fixed as a typed StoredReport, a retry narrowed to
UNIQUE-constraint failures, and a shared buildPrompt helper.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… default

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Generation takes 1-2 minutes, too long to hold a request open. Adds an
in-memory job registry (startJob/jobStatus) that HTMX polls every 2s;
generateReport commits the report to the DB before the job is marked
done, so a server restart loses job status but never a report.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…edundant loadConfig imports

renderReports had zero automated coverage for its populated/empty branches
or its HTML escaping, despite interpolating model-generated markdown into
a page. Adds tests for the empty state, the populated state (markdown,
version, template source), and escaping of HTML metacharacters — verified
to fail if escapeHtml is removed.

Also removes two redundant `await import("../config")` calls in the report
route handlers in server.ts; loadConfig is already imported statically at
the top of the file and used by every other handler.
Rename 'status' to 'rc' to avoid zsh read-only parameter conflict.
In zsh, 'status' is a read-only builtin (alias for $?). Assignment
fails silently under 'set -uo pipefail' but sets $? to 1, causing
the script to always exit 1 even on success. Renamed to 'rc' to
preserve exit code across the report call.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
- reports_dir export failures warn instead of failing an already-committed
  report, in both the CLI and the web job (FIX 1).
- introduce NoEntriesError so "no entries for the week" is a typed signal
  shared by the CLI and web handler instead of a regex on another module's
  message; the web job now finishes as a normal "nothing to report"
  outcome rather than an error (FIX 2).
- add route tests for GET /reports, GET /reports?week=, malformed-week
  400s, and GET /reports/status/:id on an unknown job (FIX 3).
- thread the generated week through the job registry so the done-branch
  reload lands back on the week that was generated (FIX 4).
- malformed week labels return a friendly 400 from both report routes
  instead of a 500, and the CLI's weekRangeForLabel call now runs inside
  the try block so it prints the intended message instead of a stack
  trace (FIX 5).
- await the two "writes nothing" rejection assertions in reports.test.ts
  so they actually check post-rejection state (FIX 6).
- split the template cache write out of the fetch's try/catch so a
  successful fetch can never fall back to a stale cache under a
  cache-write failure (FIX 7).
- restore the dropped `title` command in the CLI usage string (FIX 8).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The report was shown as escaped plain text in a <pre>, whose CSS classes did
not exist — so it rendered unstyled and without wrapping. It is markdown; show
it as markdown.

Adds a deliberately small renderer covering what reports actually use: headings,
bullets, bold, inline code, paragraphs. A full markdown library would be a new
dependency in a repo that carries almost none.

The input is model-generated text assembled from session transcripts, so the
renderer escapes FIRST and only then adds structure. Every tag it emits is one
it wrote itself; no markup can survive from the source. The ordering is the
security property, and a test pins it — including that a bold marker cannot
smuggle a tag through.

'#' demotes to h2 because the page title already owns h1.

Also adds the .report-meta and .report-markdown styles that were referenced but
never defined.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The report page was unreadable past one viewport. The app shell is fixed-height
(html, body { height: 100%; overflow: hidden }) and .full-content hides overflow
too, so anything taller than the window was clipped with no way to reach it.
The three-panel views escape this via .panel { overflow-y: auto }; the report
page had no equivalent.

Splits the page into a fixed .report-header (title, Generate button, job status)
and a scrollable .report-scroll holding the report itself, so the control stays
put while the content scrolls.

min-height: 0 on the scroller is load-bearing, not decoration: a flex child will
not shrink below its content height without it, and the container would grow
instead of scrolling.

A test pins the structural requirement — the article must sit inside the scroll
container — since the CSS alone cannot be asserted from a unit test.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant