Skip to content

feat(viz): the dashboard as the project site, with the agent's session record - #71

Merged
1nslyn merged 10 commits into
mainfrom
claude/frontend-agent-discovery-redesign-96f5cc
Sep 20, 2026
Merged

1nslyn merged 10 commits into
mainfrom
claude/frontend-agent-discovery-redesign-96f5cc

Conversation

@1nslyn

@1nslyn 1nslyn commented Sep 20, 2026

Copy link
Copy Markdown
Owner

Summary

The dashboard is rebuilt as the project site, and the agent's discovery process is opened up: every session transcript is stored in the project and shown next to the tree.

One frontend, three sources. A static single-page app (classic scripts, no build step, hash routes) reads one record layout (record/index.json, per-run graph.json, timeline.json, sessions.json, agent_links.json, nodes/<id>.json, sessions/<sid>/turns/<k>.json) from:

  • the live server, automil viz start (prints the ssh -N -L forward when it runs over SSH),
  • a snapshot written by the new automil viz export --out DIR (--single-file inlines everything into one HTML page),
  • a user's own local server, reached from the hosted site through the SSH tunnel: CORS is answered only for viz.cors_origins (default automil.org) and loopback origins, on the record routes and the event stream.

Views. The experiment tree in 3D (top-down layout on the page, kept lineage emphasised; a flat lineage tree one click away) beside the primary-value chart; a node drawer that states the keep/discard verdict with the graph's own helpers (paired or marginal SE, margin, companion guard), per-fold table, files changed, run-log tail, timing and the transcript turn that proposed the node; the discovery timeline (agent actions, queued/running bars per attempt, best value on one axis); the transcript reader (every turn, tool call, result and subagent, linked to the nodes it created; runs of tool-only turns fold into one line); nodes table; notes; runs page with the local-server connection; home and remote-access pages. Look: white page, black rules, a visible column grid on the hero, Manrope and JetBrains Mono, teal for data only (direction chosen from a four-option design canvas).

Session transcripts are a framework feature. automil activity ingest copies the Claude Code transcript and its subagent sidecar into automil/sessions/<id>/ on SessionEnd; the hidden automil activity store-sessions --root stores every journaled session (the campaign launcher's exit trap calls it; store_session_record.sh is removed). Open sessions are tailed from the runtime's own file and streamed as they grow.

Validation-only, one enforcement point. viz/record_graph.py projects every node and result through firewall.is_held_out_metric_key; no loader has a path under archive/<node>/certify/; run logs are served for terminal nodes only and re-redacted; secrets in tool results are masked. tests/viz/test_firewall_forged.py plants a test_auc everywhere and checks every payload and every opened path.

Also: the activity journal replay keeps which path finalized a session (hook-exporter-unreachable no longer replays as operator-close).

Deployment: site/vercel.json and scripts/site/publish.sh export the listed cell roots and deploy with the Vercel CLI; recorded runs stay out of git.

Test plan

  • uv run pytest tests/ green (1815 passed, 67 skipped); coverage of automil.viz + automil.session_record at 90%
  • Live server on a copied KRAS ABMIL cell in the browser: tree (3D and flat), chart, node drawer, timeline, agent transcript with folds, nodes, notes, home, runs; light and dark; phone width
  • Static export served on a second port connects to the live server (CORS + SSE live); single-file export opens
  • Forged-violation firewall tests; no test_/held_out/certify in an exported record
  • On fir: move the five finished cells' operator/session/ records into automil/sessions/<id>/ (one-liner in the campaign README) and pull the launcher change before the next discovery job
  • Create the Vercel project, attach automil.org, run scripts/site/publish.sh --deploy <cell-root>:<run-id>:<title> ...

🤖 Generated with Claude Code

1nslyn and others added 10 commits September 20, 2026 01:35
…ecord

The dashboard is now a static site that reads one record layout from three
sources: the live server next to a project, a snapshot written by
automil viz export, or a local server reached from the hosted site through
an SSH tunnel. The record is built from the project's own files
(graph.json, the per-node archive, the orchestrator log, the activity
journal) and the Claude Code session transcripts, which the framework now
stores under automil/sessions/<id>/ on SessionEnd and through
automil activity store-sessions (the launcher's exit trap uses it; the
shell script is gone).

Views: lineage tree with the value chart and a node drawer that states the
keep/discard verdict in words, the discovery timeline (agent actions,
experiment runs, best value), the transcript reader with tool calls,
results and subagents linked to the nodes they created, the nodes table,
the notes, the runs page with the local-server connection, the home page
and the remote-access page. Validation-only projection is one enforcement
point; nothing under archive/<node>/certify is ever read.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
site/vercel.json and scripts/site/publish.sh export the listed cell roots
and deploy the directory with the Vercel CLI; the record and static assets
under site/ stay out of git. The SSE stream now sends the cross-origin
grant on its own headers (a streaming response cannot take them from the
middleware), so the hosted site follows a local server live. README, the
getting-started guide and the changelog describe the new dashboard.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
_validate_journal_event stamped every attested session_end as
operator-close, so a SessionEnd hook that found the exporter already torn
down (hook-exporter-unreachable) replayed as a manual recovery. The marker
is validated against the closed set and now kept as written; a round-trip
test closes one session each way and reads both markers back.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…efault view

The pages read as a site now: full-width tinted bands, a larger type scale
on the paper palette, panes without hairline boxes, larger glyphs. The home
page draws the run's tree in three dimensions in the hero, explains the
four lanes as a strip of steps, shows the benchmark as numbers and a cell
grid, and pairs the install steps with one code panel. The tree view is
three-dimensional by default (a top-down layout on the page's paper with
matte spheres and the kept lineage in teal; the flat lineage tree is a
toggle away), the timeline lanes have backgrounds, and the value chart has
larger points.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… the transcript

Leo chose direction D from the design canvas: a white page with black
one-pixel rules and frames, a visible twelve-column grid behind the hero,
Manrope headlines, JetBrains Mono labels in capitals, teal reserved for
data, square corners. The tree stays three-dimensional (black lineage,
rimmed spheres on the page's paper, dark mode inverted). In the transcript,
consecutive turns that only call tools fold into one line that states the
turn and call counts, the tools used, the time span, failures and the nodes
created; a filter switch unfolds them.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The lane titles were drawn on the first row's line, over the first node
label and the first value tick. Each lane now reserves a header line for
its title; rows, marks and the value scale start below it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…attempt

The baseline node carries no run timing, so it was missing from the lane
and the step line began at the first improvement with no level before it.
The baseline now comes from the graph's root (dashed rule plus a label),
and the completed attempts are outlined so discarded ones no longer vanish
on the lane's background; hovering names the node and its value.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Both sides added entries under Unreleased in CHANGELOG.md; both are kept.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Main now judges the companion guard against max(quantum, k x paired SE)
and guard_basis returns that margin as a fourth value; the dashboard's
verdict unpacked three and every node export failed. The verdict payload
and the node drawer now name the guard's bar next to its delta, the same
evidence the leaderboard prints.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The shipped tcga_luad overlay now declares a submit-time policy smoke
that runs an autobench module. The framework suite copies the overlay's
registry block into its freeze fixture, and the Framework CI job installs
only automil, so the smoke could not start and the registered-policy
submission was refused on every Python version (main fails the same way
since the guard change). The fixture now carries the protected and
editable lists the freeze is made of and leaves the consumer's hook out;
the smoke validator keeps its own tests.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@1nslyn
1nslyn merged commit 6e034d1 into main Sep 20, 2026
5 checks passed
@1nslyn
1nslyn deleted the claude/frontend-agent-discovery-redesign-96f5cc branch September 20, 2026 17:49
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant