feat(viz): the dashboard as the project site, with the agent's session record - #71
Merged
Merged
Conversation
…ecord The dashboard is now a static site that reads one record layout from three sources: the live server next to a project, a snapshot written by automil viz export, or a local server reached from the hosted site through an SSH tunnel. The record is built from the project's own files (graph.json, the per-node archive, the orchestrator log, the activity journal) and the Claude Code session transcripts, which the framework now stores under automil/sessions/<id>/ on SessionEnd and through automil activity store-sessions (the launcher's exit trap uses it; the shell script is gone). Views: lineage tree with the value chart and a node drawer that states the keep/discard verdict in words, the discovery timeline (agent actions, experiment runs, best value), the transcript reader with tool calls, results and subagents linked to the nodes they created, the nodes table, the notes, the runs page with the local-server connection, the home page and the remote-access page. Validation-only projection is one enforcement point; nothing under archive/<node>/certify is ever read. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
site/vercel.json and scripts/site/publish.sh export the listed cell roots and deploy the directory with the Vercel CLI; the record and static assets under site/ stay out of git. The SSE stream now sends the cross-origin grant on its own headers (a streaming response cannot take them from the middleware), so the hosted site follows a local server live. README, the getting-started guide and the changelog describe the new dashboard. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
_validate_journal_event stamped every attested session_end as operator-close, so a SessionEnd hook that found the exporter already torn down (hook-exporter-unreachable) replayed as a manual recovery. The marker is validated against the closed set and now kept as written; a round-trip test closes one session each way and reads both markers back. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…efault view The pages read as a site now: full-width tinted bands, a larger type scale on the paper palette, panes without hairline boxes, larger glyphs. The home page draws the run's tree in three dimensions in the hero, explains the four lanes as a strip of steps, shows the benchmark as numbers and a cell grid, and pairs the install steps with one code panel. The tree view is three-dimensional by default (a top-down layout on the page's paper with matte spheres and the kept lineage in teal; the flat lineage tree is a toggle away), the timeline lanes have backgrounds, and the value chart has larger points. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… the transcript Leo chose direction D from the design canvas: a white page with black one-pixel rules and frames, a visible twelve-column grid behind the hero, Manrope headlines, JetBrains Mono labels in capitals, teal reserved for data, square corners. The tree stays three-dimensional (black lineage, rimmed spheres on the page's paper, dark mode inverted). In the transcript, consecutive turns that only call tools fold into one line that states the turn and call counts, the tools used, the time span, failures and the nodes created; a filter switch unfolds them. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The lane titles were drawn on the first row's line, over the first node label and the first value tick. Each lane now reserves a header line for its title; rows, marks and the value scale start below it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…attempt The baseline node carries no run timing, so it was missing from the lane and the step line began at the first improvement with no level before it. The baseline now comes from the graph's root (dashed rule plus a label), and the completed attempts are outlined so discarded ones no longer vanish on the lane's background; hovering names the node and its value. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Both sides added entries under Unreleased in CHANGELOG.md; both are kept. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Main now judges the companion guard against max(quantum, k x paired SE) and guard_basis returns that margin as a fourth value; the dashboard's verdict unpacked three and every node export failed. The verdict payload and the node drawer now name the guard's bar next to its delta, the same evidence the leaderboard prints. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The shipped tcga_luad overlay now declares a submit-time policy smoke that runs an autobench module. The framework suite copies the overlay's registry block into its freeze fixture, and the Framework CI job installs only automil, so the smoke could not start and the registered-policy submission was refused on every Python version (main fails the same way since the guard change). The fixture now carries the protected and editable lists the freeze is made of and leaves the consumer's hook out; the smoke validator keeps its own tests. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The dashboard is rebuilt as the project site, and the agent's discovery process is opened up: every session transcript is stored in the project and shown next to the tree.
One frontend, three sources. A static single-page app (classic scripts, no build step, hash routes) reads one record layout (
record/index.json, per-rungraph.json,timeline.json,sessions.json,agent_links.json,nodes/<id>.json,sessions/<sid>/turns/<k>.json) from:automil viz start(prints thessh -N -Lforward when it runs over SSH),automil viz export --out DIR(--single-fileinlines everything into one HTML page),viz.cors_origins(default automil.org) and loopback origins, on the record routes and the event stream.Views. The experiment tree in 3D (top-down layout on the page, kept lineage emphasised; a flat lineage tree one click away) beside the primary-value chart; a node drawer that states the keep/discard verdict with the graph's own helpers (paired or marginal SE, margin, companion guard), per-fold table, files changed, run-log tail, timing and the transcript turn that proposed the node; the discovery timeline (agent actions, queued/running bars per attempt, best value on one axis); the transcript reader (every turn, tool call, result and subagent, linked to the nodes it created; runs of tool-only turns fold into one line); nodes table; notes; runs page with the local-server connection; home and remote-access pages. Look: white page, black rules, a visible column grid on the hero, Manrope and JetBrains Mono, teal for data only (direction chosen from a four-option design canvas).
Session transcripts are a framework feature.
automil activity ingestcopies the Claude Code transcript and its subagent sidecar intoautomil/sessions/<id>/onSessionEnd; the hiddenautomil activity store-sessions --rootstores every journaled session (the campaign launcher's exit trap calls it;store_session_record.shis removed). Open sessions are tailed from the runtime's own file and streamed as they grow.Validation-only, one enforcement point.
viz/record_graph.pyprojects every node and result throughfirewall.is_held_out_metric_key; no loader has a path underarchive/<node>/certify/; run logs are served for terminal nodes only and re-redacted; secrets in tool results are masked.tests/viz/test_firewall_forged.pyplants atest_auceverywhere and checks every payload and every opened path.Also: the activity journal replay keeps which path finalized a session (
hook-exporter-unreachableno longer replays asoperator-close).Deployment:
site/vercel.jsonandscripts/site/publish.shexport the listed cell roots and deploy with the Vercel CLI; recorded runs stay out of git.Test plan
uv run pytest tests/green (1815 passed, 67 skipped); coverage ofautomil.viz+automil.session_recordat 90%test_/held_out/certifyin an exported recordoperator/session/records intoautomil/sessions/<id>/(one-liner in the campaign README) and pull the launcher change before the next discovery jobscripts/site/publish.sh --deploy <cell-root>:<run-id>:<title> ...🤖 Generated with Claude Code