Repository navigation
Conversation
`fly`: a Drosophila brain, three assays, an HTTP backend and the page a person watches it learn on. The four domains before it grow a library; this one is the first asked for a frontend, so the specification alone has to carry the design. It is also the first shipped with no reference implementation, and the cost is stated rather than hidden: no offline actor, so `--domain fly` without a model now says so instead of raising AttributeError, and no kill_report, so nothing in the repository checks what this suite rejects. What stands in its place is that every number in the specification was measured on a throwaway implementation while the suite was being written -- 91 assertions passing, and three of the specification's least obvious paragraphs are there because that implementation failed them first: a ring attractor sharpened with any pointwise nonlinearity quantises heading to the nearest wedge; an avoidance assay with the punished source in the middle of a walled arena measures the arena, not the animal; and `src/learn/` beside a public `learn()` is two things called `src.learn`, where the submodule import rebinds the name. That last one is the same collision this port was bitten by twice from the other direction, which is why the specification names it and test_integration.py asserts it. 93 tasks over 8 test files, 10 of them an audit set no agent can read. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
…e suite Upstream the agents write the tests. They have file and shell tools, they write `test/*.exs`, and "tests are the definition of done" (agents/manager.ex); the external suites Genesis validates against -- c-testsuite, LLVM, Csmith -- are the experimenter's measurement, applied afterwards. This port had it backwards. A human wrote every assertion, and then `TEST_FAILURE` pasted its **source** into the prompt, which is the strongest hint there is: an agent handed the assertion is not implementing a specification, it is writing to an assertion. `TestSuite.hidden` is the other way round. The acceptance suite never enters the repository -- it is merged into the scratch copy one evaluation sees, the way the audit set already was -- and the prompt becomes the requirement it was written from plus the assertion's name. The name carries real information and that is deliberate: `test_the_code_is_sparse_at_every_concentration` is a requirement written as a sentence. The threshold, the tolerance and the inputs are exactly what it does not carry. So every test *in* the repository is an agent's own, and two new things can be asked of a child. `own_test_failures` runs them in one interpreter -- `mix test`, theirs -- scoped to one node's subtree. `own_review` refuses a child whose own tests fail, and a child that wrote an implementation file and no test for it: upstream every node is accountable for its own subtree, so a node with untested code is a node whose parent has nothing to run. Also `ConcurrencyGauge`, because the obvious arithmetic is wrong. `usage.seconds / wallclock` looks like it says how many calls were in flight and does not -- `seconds` spans phases that run before `evolve()` does, while a stage profile's `wallclock` covers only the stage, and the md run's ratio came out at 8.2 with four workers. It counts them directly instead, and the run reports the peak. In `examples/`, not in the engine: what the engine's concurrency *is* is its own business; what an example observed is the example's. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
…he rest The domain now hands over two things and both are requirements rather than design. `REQUIREMENTS.md` is what the user asked for, in their words -- a connectome-grounded Drosophila brain, three assays, learning that visibly changes behaviour, a frontend and a backend, and a test beside every directory's code -- plus a reading list. No equations, no coefficients, no module layout. `fly.py` is that requirement made executable, and it imports exactly four names from `src`, so the public surface is four functions and everything else is the agents' internal business. Everything else grows: the CONTEXT.md tree, the circuit, the layout, the frontend, and every test in the repository. Both assertion sets stay outside it. `hidden` is black-box acceptance and drives the search -- the agent is shown the requirement heading, the assertion's name as a sentence, and what it reported, never a line of source. `audit` is sealed: written before the run, never in the repository, the same requirements asked with different seeds and different tasks, and run once at the end. The gap between "acceptance passes" and "the sealed suite passes" is the measurement of whether it built the thing or wrote to the acceptance. `suite_failures` now asks what upstream's manager asks -- do *your own* tests pass -- because the human suite is not something this domain's agents can see. `own_review` joins the review chain, so a child that wrote an implementation file and no test for it is sent back. One requirement neither set can check, stated rather than hidden: whether the brain is really built from the connectome. Black-box cannot ask. Only the parent reading the diff, and a person reading the result, can. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
…inary The SDK path needs an `ANTHROPIC_API_KEY`. The CLI is authenticated another way entirely -- an OAuth session, a subscription, a corporate login -- and on a machine where that is the credential that exists there was no way to run a port at all. This session hit exactly that: the endpoint returned `401 ModelArts.81015, Invalid client ip ... rejected by api key ip whitelist setting`, and `preflight` caught it in one call rather than in 400 episodes of silent nothing. `--provider claude-cli` is a one-shot `claude -p` with every tool denied, which is a plain prompt-to-text function and so is exactly what a Completion is. Two caveats are in the docstring rather than hidden, because both change the output and not only the plumbing: the CLI wraps the prompt in its own system prompt, tool definitions and any CLAUDE.md it finds, so a call carries tens of thousands of cached input tokens the API path would not; and the cost lands on the CLI's credentials, where no budget flag here can see it. Token accounting counts cache reads and writes as the prompt tokens they are -- `input_tokens` alone reported 10 for a call that really carried 28,000. And `--no-thinking` says it has no effect here instead of being dropped in silence. Two fixes the same run needed. The Claude Code executor now runs the model the run asked for, so the architect and every executor episode are not on different models by default. And its failure template comes from the domain rather than a constant: handing a blind domain's session the assertion's source is the one thing that domain exists to prevent. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
… budget Three of the four gaps the deviations table recorded, closed. **Upstream's architect Phase 3.** It does not stop at design: it reviews the implementation and re-spawns refinement architects where a node misaligns (`agents/architect.ex`). This port stopped after the design, so a record written before any code existed stayed the map for ever. `--refine` checks a node's record against what is actually in the directory -- a routing table promising a child nobody created, an `## API Surface` that never mentions a file sitting right there -- and where it has drifted, re-spawns an architect on that one node to rewrite it against the code as it is. **And that is also the 62 updates.** Upstream's archive shows 26 `CONTEXT.md` creations and 62 later accepted updates affecting 19 files: a record is maintained, not written once. This port wrote them on two occasions, and a `--mode a` run sat at 0.938 for 30 003 rollouts reading a map of a layout the work had already left behind, with nothing in the mechanism able to say so. The hook fires on the leaf path as well as the accountability pass, and leaves need it more -- a leaf is where the code lands, so its API Surface is the first thing to go stale. **One episode, two budgets.** Upstream gives a root agent up to 2,048 model-tool turns and a child 128. The Claude Code executor had one number, 24, for both, which either starves the root or hands every leaf a session it has no use for. It now carries upstream's own two numbers, and a single `--max-turns` still means a single number for a cheap run. **And the node budget.** The fly domain's architect named 17 children it never reached at the default of 12, so the tree came out two deep and truncated. `--nodes` sets it. The fourth gap, the human merge on the dashboard, stays recorded and unbuilt. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
…ing it The executor learned upstream's two budgets -- 2048 at the root, 128 below -- and then the parser's own default of 24 fed straight past them, so a run configured for the split reported "up to 24 turns" and meant it. The default is now 0, which means "use the split", and a number given explicitly still pins both to it for a cheap run. The header prints both numbers rather than one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
The fly domain's architect spent 27 minutes designing 32 nodes and was about to hit `--nodes 40` with a queue still behind it. Raising the budget meant paying for all 32 again, because `design()` asked for every node it reached whether or not a record was already there. With `resume`, a node that already carries a record is kept: its routing table is taken as the design, its children are queued, and no call is spent. The thing that revises a record once code exists is the refinement architect, not this. Off by default, and deliberately. At `root_path=""` the record already in `given` is the *harness* record -- generated, not designed -- so reusing it would skip the root and leave the tree with no design at all. The runner turns it on only for `--continue-from`, where a record that is there really is an earlier phase 1's work. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
The routing line's right-hand side is not decoration -- it is the objective the parent's architect handed that child. A phase 1 continuing from records rather than from replies has nothing else to give one, and the first version passed down the *parent's* objective instead. What that does is visible in the run that found it. An architect asked to design `src/brain/circuits/cell_types/mushroom_body`, while carrying the root objective "implement the software REQUIREMENTS.md asks for", came back with children `brain, learning, environment, simulation` -- it had redesigned the whole library from the top, at depth five, three nodes running. `parse_routes` returns `(path, what it handles)` where `parse_routing` returns the paths alone, and resume queues each child with its own line. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
…ctions Upstream calls the routing table "your primary delegation tool ... the map that makes recursive delegation work". A record whose table disagrees with the children the architect just opened is a broken map, and phase 1 now repairs it rather than passing it on. Dropping a refused entry was already there -- a table advertising a node nobody may write is a trap for the next manager, which would delegate there and be refused in turn. Adding a forgotten one turned out to matter more. An architect answered with five children and a `## Routing Table` section holding its API Surface instead: "`__init__.py` — Exports get_regions(...)", then two sentences about FlyWire. Nothing looked wrong that run, because the children were queued from the *reply*. Then the budget ran out, a later phase 1 resumed from records, rebuilt its queue from the **table**, found no routes in it, and dropped an entire subtree of the brain -- five regions, silently, because the record and the reply had disagreed and only the reply was ever right. And a second guard from the same run. "If the objective feels too large, that is exactly the signal to decompose MORE aggressively" has a counterweight upstream states in the same breath: single responsibility, and shared capability belongs at the lowest common ancestor. A node named after one of its own ancestors is that rule broken in the one way a tree can show -- what is really in there belongs to the ancestor, or the ancestor's name was wrong, and either way two places claim it and no later agent can tell which. The run produced `.../receptor/types/catalog/types` and `.../catalog/demographics` beside an existing `.../receptor/population/demographics`. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
… at a time **The driver was a choice, and this port's own deviations table said so**: "a frozen entry point, so the result is a program rather than a package nobody can invoke. A choice, and a defensible one." Upstream Mode B takes an objective and nothing else; c-testsuite, LLVM and Csmith are the experimenter's benchmark, applied afterwards, never in the repository. Upstream does have a contract -- it is just implicit in the objective. "Build a C compiler" pins `cc -o foo foo.c`, an convention everyone already knows, which is why c-testsuite can run at all. "A fruit fly simulator" has no such convention, so it has to be said -- but it belongs in the *requirement* (how I want to use it) rather than in the repository (here is code you may not touch). `REQUIREMENTS.md` §4 now says which three subcommands must work and that each takes `--format json`, because a person who wants to plot a learning curve asks for machine-readable output. `fly.py` is theirs to write. So both suites are black box now. They start `python fly.py ...`, read its JSON, drive its HTTP server over a socket, and import nothing -- which is the shape of upstream's own validation. One file is frozen: the requirement. **And phase 1 designs a level at a time.** Upstream *spawns* sub-architects, which is a statement about independence: every node at a depth inherits the chain down to its own parent, designed a level ago, so nothing in a level can depend on anything else in it. Running them one after another was this port's choice and it cost the fly domain 32 minutes for 71 nodes. The asks read a snapshot rather than the live tree, so a sibling can never see another sibling's record even if it finishes first, and the tree does not depend on which call returns when. Two bugs the change surfaced, both fixed: `_apply` still carried the serial loop's own `self._ask`, so every node was asked twice; and resume was decided after the dispatch, which spends exactly the call it exists to save. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
A Claude Code episode goes through the local CLI whatever `--provider` says, and the CLI was handed no model at all unless the provider happened to be `claude-cli` -- so a run whose architect was on an API model had every executor session on the CLI's default instead, two models, in silence. It now takes the run's model, and `--executor-model` is there for the case where they should deliberately differ. The CLI reads ANTHROPIC_BASE_URL and ANTHROPIC_API_KEY from the environment, so pointing the SDK and the CLI at one endpoint is all it takes to put both halves of a run on the same model. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
It printed nothing until the whole phase finished. On the fly domain that was 82 minutes of silence, and answering "how far in is it" meant parsing the Claude CLI's own session log out of ~/.claude/projects. Then the architect moved to the SDK, the subprocess and its log went away, and the question had no answer at all -- six sockets open and a thread count. Each node now prints as it lands, with the running count against the budget. `--quiet-phase1` turns it off. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
Upstream says both things in the same breath -- decompose MORE aggressively when the objective feels large, *and* single responsibility, shared capability at the lowest common ancestor, a directory holding one short function is a directory that did not want splitting. The first is a command. The second needs judgement, and a model handed both executes the first. Measured: the fly domain's architect designed 600 nodes for one simulator -- 921 before the budget cut it -- three quarters of them pure routing, at depth 8. It gave `apl`, which is a single GABAergic neuron, two children and one of those three more. Upstream's 123-hour C compiler run, 750 files and 249 000 physical lines, has **26** nodes and bottomed out at depth 5. So the counterweight is enforced rather than stated. Each child now declares `files` -- how many source files the architect expects that directory to hold, counting everything below it -- and a child of two or fewer is refused: it is two files in this node's API Surface, not a directory. A node of three or fewer is told to return no children at all. An undeclared count is not a refusal; the guard reads what is there rather than inventing a number. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
`_parse` now returns `files` on every child, undeclared as 0, and the assertion that pinned the child dict verbatim had to say so. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
`--executor claude-code` spawned the CLI with an inherited environment. Inside a managed Claude Code session that environment carries CLAUDE_CODE_REMOTE, which puts the CLI on the host's session ingress and makes it ignore the ANTHROPIC_BASE_URL and key it was handed. So every episode sent the host's token to a third-party endpoint and got 401: one fly run opened 28 sessions, failed 23, and wrote no files at all. What the CLI reports for this is "Authentication error - This may be a temporary network issue, please try again", which is neither, and the stderr beside it names an unrecognized model for query_source generate_session_title -- the session-title side query, not the agent turn, and printed just the same on runs that work. I read that line as the cause and was wrong. `cli_env` drops the host-provider variables, but only when the run actually points somewhere else; a session on the host's own provider is untouched. The credential still only ever travels from the environment into a child's env. This also corrects 611c4af, which claimed pointing the SDK and the CLI at one endpoint was all it took to put both halves of a run on the same model. It is what the CLI documents, and it is true on a laptop. It was not true here, and the port's own fly runs are what proved otherwise. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
`own_test_failures` returns [] when the repository holds no test files, which reads as "everything passes" -- and that answer is the precondition for `complete_task`. A fly root agent used it: sixteen implementation files, zero tests, acceptance reward 0.000, and it claimed the objective delivered five times and was believed every time. That is backwards from the rule the same class already enforces on every child in `own_review`: code with no test is not finished. `require_tests` applies it to the whole repository, and the blind domain's `suite_failures` passes it, because that path is a gate. It stays off where the question really is "did anything the agents wrote break" -- a parent reviewing a child that has legitimately returned no code yet should not be told its tests failed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
The architect asked for a repository design and got 1708 characters of a JSON
object that stopped mid-word. The run reported "architect designed 0 nodes,
1 replies unusable" and then grew a whole phase 2 with no tree at all, which
is a shape that looks like a bad model or a parser bug and is neither.
The cap was the default 4096, and this backend spends it on thinking first,
so what was left for visible content did not hold the record. The same file
already warns about this ("a 1024 cap returned nothing at all for 4 of 8
reflection prompts") -- it just had no way to raise the number from a command
line. Turning thinking off is the other lever and the worse one: it changes
what the model produces, not only how much. Measured on the architect prompt:
4096 emits ~15k output tokens across attempts in 148s and truncates about half
the time; 16000 takes 264s and parses.
`--max-tokens` and `--timeout` are declared and honoured here, which is what
this module exists for. They come in together because raising one without the
other only moves the failure: the call that now fits took 264s, and claude()'s
120s default would have timed it out three times over.
Both are opt-in, by naming a default. Nine ports already declare `--max-tokens`
themselves, each with a number it measured -- 4096, 16000, 32000 -- so
declaring it here unconditionally is an argparse conflict that takes those
entry points down at import, which is how I found them: six tests across
porous, openevolve and metasearch. A port that wants the shared flag asks for
it, the way `include_val_cap` withholds one from a port whose splits are
already frozen. The method runner asks, and keeps its measured 1024/180.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
Upstream's Architect is not a question with an answer. `agents/architect.ex` is
`use EvoGit.Agent` -- "an agent session loop template that manages a single
agent session, handling tool loops" -- with `agent_type :read_write`, which in
`agent/tools.ex` means file_read/create/write/edit, make_dir, context_read/
write/edit, run_bash, ripgrep, glob, list_dir. CONTEXT.md is written with a
tool. Nothing comes back as a structured reply and there is nothing to parse.
So is the Executor, and so is the Manager: every agent there is a session.
This port had the Executor right already (`--executor claude-code`) and the
architect as one completion per node returning {"record", "children"}. That is
the shape a reasoning model is slowest at -- one reply holding a whole record
and a whole child list, thought about once, at length -- and a truncated reply
is not a short record but an empty phase 1: one run reported "architect
designed 0 nodes, 1 replies unusable" and grew a phase 2 against no tree.
`--architect-session` designs a node the way the executor implements one: a
session in a throwaway copy, tools Read/Write/Edit/Glob/Grep, and what it did
recovered by reading the file it wrote. No Bash -- the executor gets a shell
because it must run the suite it is judged by, and an architect that can run
things is an architect that starts implementing.
One node per session, with the phase still driving the levels. Upstream's
architect spawns sub-architects and recurses inside the session; keeping the
recursion in the driver keeps phase 1's level-parallelism, keeps a session
small enough to watch, and keeps the spatial contract checkable per node.
That is a stated difference, not an oversight.
Two things fall out of the record being a file:
The design rules now live in ARCHITECT_RULES, and each path appends its own
delivery -- "reply with JSON" or "write the file". Duplicating them is how the
two paths would grow different trees.
A routing line may declare its child's size, `(12 files)`, because the
file-count threshold that refuses a child too small to be a directory has
nowhere else to live once there is no JSON. Records without counts still route.
And one defect this found, which was costing the completion path too: the
routing section was matched by an exact prefix on "## Routing Table", so
"### Routing Table", "## Routing" and "## 4. Routing Table" all opened nothing
and the node silently became a leaf. A session wrote a complete five-child
table under such a heading and the tree recorded zero children.
Measured on one fly node against a coding-plan endpoint: session 497s
uncapped, 203-242s with a 2048-token reasoning cap, same 4.6k record and the
same children. `--thinking-tokens` exposes that cap, following upstream, which
carries `reasoning_effort` per model profile beside `max_tokens` and
`concurrency` rather than fixing it. It bounds reasoning; it does not disable
it, which changes what the model produces and not only how long it takes.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
`agents/manager.ex` opens with "The Manager does NOT implement features directly" and lists five jobs: analyse, plan, delegate, **validate results**, report completion. `agents/context_extractor.ex` is Mode A's root agent and the reason Mode A exists -- a repository with code and no CONTEXT.md cannot be worked on by recursive delegation, because the routing table is the map. Both are sessions with tools upstream. Both were one completion returning JSON here. Two of the Manager's five jobs are model calls in this port, and each now has a session behind it. Planning and delegation. The completion decided from the routing table and the record alone. The routing table exists so a parent *need not* investigate the subtree -- upstream says so -- but need not is not cannot, and the Manager is the role upstream gives read tools to precisely so it can check. Validating results. This is the one that was actually wrong. The completion was handed a diff rendered and truncated at 12 000 characters; a reviewer judging a truncated rendering of the work is not reviewing the work. `ReviewSession` runs over the parent's state with the child's edits applied, so Read, Glob and Grep reach the files the child really wrote. The extractor's gap was already written down in `_extract._ask`: "Upstream it reads what it needs with a tool; an agent here has none, so the node's own files travel in the prompt" -- eight files of six thousand characters, and the run before that cap invented an API surface off the file names. It reads the code now. Its record can also carry the sections upstream lists for it and an architect has no use for: Design Decisions, Notes for Agents, Dependencies, Test Strategy, on upstream's own rule that a section earns its place if it "would save an agent from re-investigating or re-discovering something". The session machinery moves to `_session.AgentSession` and the architect sits on it, so there is one copy of materialize / run / read-back rather than two that drift. Tool sets follow `agent/tools.ex`: `:read` roles get Read/Glob/Grep plus the one write a read-only agent does make (its own record), `:read_write` adds Edit. Nobody but the executor gets Bash -- the implementer has to run the suite it is judged by, and a design or review session that can run things is one that starts implementing. Deliverables are files, and lines rather than JSON: a plan of five children whose fourth line is malformed should delegate four, not none. A verdict that does not say ACCEPT or REJECT is not a rejection -- the child did the work, and a reviewer that cannot speak is not evidence against it. Still this port's standing difference, stated rather than hidden: the recursion lives in the driver. Upstream's agents spawn subagents and recurse inside the session; here one session does one node and the phase or the delegation policy drives the next. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
A fly formation run reported `sessions=52 failed=43` -- 83% of implementation episodes judged failed -- while the parent review rejected only three and the spatial contract dropped nothing. The CLI writes a transcript per session, and all 52 were on disk, so the cause was measurable rather than inferable: 27 ran into the wall (>=860s of a 900s limit), 25 killed mid-tool-call 17 had the endpoint drop the stream first (26s-830s, also mid-tool) 8 ended on a text turn, having finished 0 reached the 128-turn child budget -- the busiest made 97 Three defects, not a tuning problem: * The executor was the one role built without the run's session settings. Every other role got `**_session_kwargs()`; this one got three arguments of its own, so `--timeout 600` and `--thinking-tokens 2048` reached the architect, the manager, the reviewer and the extractor and not the role that writes the code. It ran the whole domain at a 900s default nobody chose, with no reasoning cap. * `--timeout` is the timeout on one model call (`_common` says so in its own help, and genesis defaults it to 120s). A session is a loop of many calls, so the wall is now its own flag, `--session-timeout`, printed in the run header beside the turn budget. * `turns=309` was not the truth. There is no JSON after a SIGKILL, so `num_turns` is never read for a session that hit the wall; the transcripts hold 2607 assistant turns. Timeouts are now counted apart from every other failure (`failed=43 timeout=27`), because the two have different fixes. Work was never discarded: the port reads the worktree, not the exit status, so a session interrupted at the wall still returns what it wrote. That was already right and is now covered by a test. Also, from the same transcripts: `subprocess.run(timeout=)` signals the CLI and nothing else. A session is told to run the suite, a suite run starts servers, and one was still listening two hours later with its working directory deleted. Sessions now lead their own process group, killed on the way out -- after the wall and after a clean finish alike. Two smaller things found while reading: `--mode a --agent-sessions` raised NameError, because the extractor read `use_sessions` and `_session_kwargs` from above their definitions; and `ArchitectSession.strays` reported 0 unconditionally from a helper that walked the workspace with a module its file never imported. The session owns the workspace and deletes it, so it does the counting now. Full suite green; six new tests, including one that asserts a session's grandchild does not outlive the episode that started it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
Reading the executor transcripts for the wall turned up what else they carried. A run launched from inside a Claude Code session had every episode inheriting that session's whole situation: its MCP servers, its skill listing, its agent types, the user's email address, and a system prompt about reviewing pull requests and publishing artifacts. Same endpoint, same one-line prompt: inherited 31,850 input tokens 6.7s 224 KB transcript --bare --strict-mcp-config 1,317 input tokens 2.4s 16 KB transcript A 24x prefix on every turn of every episode, none of it about the objective and some of it competing with it -- an executor told it is accountable for a pull request has been given a second job. It also feeds the wall: 30k extra tokens per turn, ~50 turns a session, 52 sessions. Three things were shared and are now not: * Identity. CLAUDE_CODE_SESSION_ID and its siblings are dropped, so each episode is its own session. Inherited, all 52 wrote transcripts named with the host's session id, and their TodoWrite state -- keyed by that id -- landed in the host's own task list, two hundred entries of "Implement src/brain package". * State. CLAUDE_CONFIG_DIR points at one directory per run, so transcripts, todos and synced skills stay out of ~/.claude, which one run had left 685 project directories in. Reported at the end of the run and deliberately not deleted: those transcripts are the only record of what an episode did, and reading 52 of them is how the wall was found. * Context. --bare --strict-mcp-config, dropping hooks, LSP, plugin sync, commit attribution, auto-memory, MCP servers and CLAUDE.md auto-discovery. The artifact carries CONTEXT.md records the brief names, not a CLAUDE.md. Bare mode reads credentials strictly from ANTHROPIC_API_KEY, never OAuth and never the keychain, so it is used only when the run brought its own key: a session billed to the local CLI's sign-in makes no API call at all with it (measured, duration_api_ms 0), which is a worse failure than a long prompt. The three fences are untouched. A bare session still honours .claude/settings.local.json -- checked against the real CLI by asking one to append to a denied path and watching it refuse, while it read and wrote the files it was entitled to. Full suite green (2939 tests). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
The previous commit described it as "a system prompt about reviewing pull requests and publishing artifacts" and an executor "given a second job". That was wrong, and the transcript's own prompt_snapshot says so: the system prompt is 5,720 characters, and the tool schemas are 178,742. 26 of them, of which Artifact alone is 64,168, then Monitor at 14,335 and DesignSync at 13,255. The executor is allowed six tools; Read and Bash together are 6,887 characters of that list. Beside it ride a 13.5 KB skill listing, a 3 KB agent listing, 900 bytes of deferred tool names, and the user's email. And one measurement that was missing, which is the whole mechanism: inherited 31,850 input tokens with the port's --allowedTools/--disallowedTools 31,099 --bare --strict-mcp-config 1,317 Permission flags do not shorten the request. They say what the session may *call*; every other schema is sent regardless, and the port had been passing those flags all along. Nobody instructed the episode to review a pull request -- it was handed the equipment of the session that launched it, and equipment is priced per turn whether or not it is reachable. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
…e fences are Reading one executor's full trajectory from the isolated run turned up two things the code was wrong about. **--bare does not expose Write.** Neither --tools nor --allowedTools brings it back: a bare session given `--tools Read,Write,Glob,Grep` answers "I only have a file Read tool available". One executor spent a turn finding out -- `No such tool available: Write. Write is disabled for this session` -- and then wrote every file through `cat > f << EOF`, while the same code without --bare had made 1,910 successful Write calls. Edit is there and creates a file that does not exist, which is how all seven phase-1 records in that run were written by sessions with no shell at all. So the tool list now drops Write when bare is on and keeps a write path that exists. **--allowedTools is not a fence.** It is the auto-approve list; under --permission-mode acceptEdits a session reaches for whatever built-in tool it likes. Measured over that run's role sessions: the manager used Edit 41 times and the reviewer 6, and Edit was in neither one's allowed tools. What does hold is --disallowedTools -- no role session ran a shell or reached the network -- and AgentSession.run, which returns the paths the caller named and drops everything else. The comments and the test that claimed otherwise now say this, and the test asserts the fences that exist rather than the one that does not. **The worktree bounds what survives, not what is seen.** Four of twelve episodes ran `find /` and read a previous run's output from /tmp; one opened the very _cli.py that answered the acceptance failure it had been handed to reproduce. The brief now forbids reading outside the checkout, which is a rule and not a wall: enforcing it takes a container, and short of that a machine should be cleaned of earlier runs' output before starting one. Said plainly in the module header and in the doc rather than left as an implied guarantee. Full suite green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
The blind property was a line in a prompt. Four of twelve episodes in one run
ran `find /` and opened a previous run's output from /tmp; one of them read the
very _cli.py that answered the acceptance failure it had been handed to
reproduce. The worktree bounds what a session's work *becomes* -- an edit
outside the node is a request, a frozen write is dropped, nothing survives but
the diff -- and none of that bounds what it can read.
agentdescent already owns the fix and I had missed it. sandbox.py manages
*lifetime* and says so ("It does not manage isolation"); sandbox_container.py is
titled "A sandbox that is actually a boundary" and gives exactly what is wanted:
only the workspace visible, read-only root, no capabilities, no new privileges,
resource ceilings, network off unless the spec asks. The first look here
reported no engine and that was wrong -- dockerd was installed and simply not
running.
examples/genesis/_sandbox.py subclasses ContainerProvider and adds the mounts an
*agent* session needs that a candidate's test run does not: the claude binary's
own install (the image has no agent in it), the proxy's CA bundle (a session
that cannot verify a TLS-inspecting proxy spends its turns on certificate errors
-- a probe lost two to pip's CERTIFICATE_VERIFY_FAILED), and a per-session CLI
state directory so the transcript outlives the container. Nothing else of the
host. Asked from inside, on this machine:
ls /home/user/agentdescent No such file or directory
find / -name 'algo-genesis.md' (nothing)
ls /tmp | wc -l 0
ls /work CONTEXT.md md.py spec src tests
touch /etc/x Read-only file system
grep CapEff /proc/self/status CapEff: 0000000000000000
Four things this needed that the flags alone did not give:
* `network="inherit"` leaves the engine's default bridge, which is not the
host's network. The endpoint here is reached through a proxy on the host's
loopback, and on a bridge 127.0.0.1 is the container. The subclass adds
--network host.
* The binary under the mount is the *resolved* path: `claude` is a symlink out
of a node install, and exec'ing the link's own path inside is "stat: no such
file or directory".
* Only the agent's argv[0] is rewritten. A shell probe through the same
workspace is not the agent, and rewriting it turns `sh -c 'ls /'` into
"Please run /login".
* The provider's own .agentdescent-mount marker is not the session's work; the
first sandboxed diff reported it as an edit the node never made.
Also dropped MAX_THINKING_TOKENS from the inherited environment. It is not
identity but it leaks the same way: the reasoning cap the *host* session runs
under silently overrode --thinking-tokens.
Two honest limits, both stated in the doc: the network is on, so this is a
boundary against contamination and not against hostile code; and where no engine
answers the run says so once and falls back to a plain directory rather than
pretending.
Full suite green.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
The default image is a name rather than a build, so a run never fails at its first episode for want of a build context. What it does not carry is a test runner -- and an episode is told to run the suite it is judged by, so every one of them pays for installing it: a probe session spent two of its turns on pip, one of them on the proxy's CERTIFICATE_VERIFY_FAILED. --sandbox-image lets a machine that has built one point at it. Two lines and a default that does not change. Genesis tests green (the driver's parser is exercised by test_the_session_wall_is_not_the_model_call_timeout); the full suite was green on the commit before this one and this changes one flag. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
The driver hands every session role `_session_kwargs()` -- model, wall, reasoning cap, and now the sandbox. Four of the five take `**kwargs` into AgentSession and picked it up; ArchitectSession spells its signature out, so adding the sandbox silently dropped it out of that contract. The run died at phase 1 with `unexpected keyword argument 'sandbox'` *after* printing its whole header, which is the worst shape for this failure: eighteen lines of a working run, then a traceback, and a background task that reports exit 0 a minute after it started. A test now builds all five roles from one kwarg set and asserts each of them carries it, which is the contract the driver has always assumed and never checked. Full suite green. The sandboxed run is past its first phase-1 node with six containers up. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
… next start A session container outlives its `release` when the process holding it dies. One run was four minutes into phase 2 with eight episodes in flight when the machine running it restarted, and afterwards all eight were still up and idling, holding memory and a workspace mount apiece. The engine already knows how to find them -- they carry its label and a start time, and `ContainerProvider.reap` removes the ones past the TTL, which by construction belong to nobody. Calling it once when the sandbox is built means the next run cleans up after the last one, and the count goes in the run's own report rather than being something you have to go and look for. Housekeeping never fails a run: a reap that raises is a reap that returns zero. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
An episode is told to run the suite it is judged by. A suite run leaves a `.pytest_cache`, and `_read_tree` skipped `.git`, `.claude` and `__pycache__` but not that -- so the directory reached the accepted version and was materialised into every later workspace, where the next agent reads it as if it were the project. Caught four hours into a live run, in the transcripts. A manager session situated at `src/arena` opened `src/arena/.pytest_cache/`, spent seven turns reading `.gitignore`, `CACHEDIR.TAG`, `README.md` and `v/cache/lastfailed`, and then wrote its plan to `src/arena/.pytest_cache/.genesis/plan.md`. AgentSession reads back the path the caller named and nothing else, so that plan was never read: the episode cost a whole session and returned nothing. One workspace had nested the directory twice -- `src/brain/central_complex/.pytest_cache/.pytest_cache/`. It also shows up from outside as the thing that made the run look like it was slowing down: new code files per check went 18, 27, 14, 10, 4 while sessions per check went 21, 33, 33, 36, 54. ARTIFACT_DIRS now names what a session *makes* rather than writes, and both readers skip it -- the executor's diff and the role sessions' stray count. A virtual environment is in the list without having been seen yet: it is the same failure with three orders of magnitude more files. Full suite green. The live run started before this landed and does not get it; the next one does. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
ARTIFACT_DIRS keeps `.pytest_cache` out of the accepted version, which is where the problem starts. It is not where it ends. In the run that turned this up the directory got in anyway, and the managers delegated *into* it: 21 executor episodes were situated at a cache directory, nested as deep as `src/brain/central_complex/.pytest_cache/.pytest_cache/.pytest_cache`. Each was an episode, a container and a session spent reading `CACHEDIR.TAG`. So the delegation refuses it too. A node is a *source* directory; a dot directory is never one, and neither is anything else in ARTIFACT_DIRS. Counted as a mistaken node, which is exactly what it is -- a path read as a node when it is not one -- beside the existing refusals for a file and for a module the path would shadow. Two fences on one failure because they fail differently: the first stops the directory existing in the state at all, and the second holds even when something else puts one there. Full suite green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
A run's phase 1 reported `architect designed 1 nodes, deepest 0` from a record that named six children in plain sight. The architect session was fine -- six turns, no errors, a 3,924-character record with a full routing table. The parser threw it away. The brief asks for `(12 files)` and the pattern read `\d+`, so `(~32 files)` -- "about thirty-two", a perfectly reasonable thing to write -- did not match. The count sat in an optional group that still had to match wherever it appeared, so failing it failed the *whole line*: the child did not lose its size, it vanished. Six lines, six vanished children, one node. The bracket now swallows anything and the number is dug out of it afterwards, so `(~32 files)`, `(about 32 files)`, `(32 files, maybe more)` and `(several files)` all keep their child, and only the last loses its count. Rare and expensive, which is the worst combination: across every run on this machine, 7,244 routing lines carried an exact count and 49 an approximate one -- 0.7%, and one of those 49 landed on a root architect and cost a whole run. A one-node tree has no delegation, no parent review and no spatial contract in it at all. Full suite green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
The blind failure ended "Write your own test that reproduces this, put it beside the code it covers, and make both pass". Every agent in a rollout is shown that same text, and at a node it reads as an order to fix the import here -- so eight different nodes each wrote their own _cli.py, 79 writes between them, src/brain/olfactory/_cli.py alone 21 times. Only the one at the repository root could ever have resolved the root import; the rest were local imitations of a file that has to exist elsewhere, and work the parent then had to undo. The evidence now says it is the repository's, not necessarily yours: judge it against what you own, write the test only if the fix belongs under your path, and otherwise name the path in the final message and leave it alone. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
The spatial contract caps a proposal at max_files_per_diff, and the cap rejects the diff rather than the files over the line -- so an episode one file past it contributes nothing. At 6, counting distinct paths written across 309 productive sessions, 11% of the episodes that did work lost all of it; the largest was carrying 23 files. Nothing said so: the two existing counters are about authority (wrote outside its subtree, mistook a file for a node) and neither moves for this, so a run that lost an eighth of its work and a run whose agents had nothing to say printed the same header. Two fixes. The drop is counted, as discarded_diffs=N (M files) in the world summary. And the cap stops pretending to be a trust region: the trust region is the node's subtree, enforced edit by edit, and twelve files under one node are not more dangerous than six. For a session executor it is now a runaway guard at 64 -- three times the largest legitimate episode, where only a loop that dumps a tree reaches it. 24 is where it stops binding at all. A single completion asked for whole files keeps 6, because there six really is one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
Between "Growing the world" and the final summary the run printed nothing at all, and for a formation run that is hours. Asking whether anything was being accepted meant reading CLI transcripts and looking inside live containers -- and a live worktree holds the base state PLUS whatever the session has written since, which is not the accepted state. Two readings out of three drawn that way were wrong. evolve() already takes an on_round hook; the run now passes one and prints a line per merger sweep with reward, artifact size, committed, rejected, rollouts and elapsed. RoundInfo.reasons goes on the end because committed=0 has two causes that need opposite fixes: the gate refused the work, or the work never reached the gate. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
A rollout is a walk of the whole Context Tree with one Claude Code session per node, and it was a for loop. On the fly tree -- 8 nodes, deepest 2 -- a session takes five to fifteen minutes, so a rollout took about an hour. py-spy on a live run, 75 minutes in, showed all four workers still inside their FIRST propose(), three levels deep in nested _episode frames. The accepted state held three files. That read as an acceptance problem and was not one: the parent gate accepts a tie, four of the five reviews that had run said ACCEPT, and the staleness policy keeps a card whose reward is merely unchanged. Nothing was being rejected. Nothing had finished. And --workers counts whole tree walks, so four workers were four independent descents of the same eight nodes -- two of them on src/brain/navigation at the same moment, doing the same node's work, of which one result could survive. Siblings now run concurrently and are folded in delegations order, so the proposal a rollout returns is unchanged. The spatial contract is what makes that safe: a child writes only under its own subtree, so two siblings cannot touch the same path. --node-workers says how wide one level may get, --max-sessions bounds what the endpoint actually sees, and the bound lives in run_cli rather than in the pools -- sessions are the scarce thing, and they never nest, so one semaphore there covers the whole tree with no way for a parent to deadlock on its own children. Every counter the port reports is now incremented under a lock. x += 1 is a read and a write with a bytecode boundary between them, and these numbers are the port's evidence. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
… of them --bare removes the inherited tool schemas and a signed-in run cannot use it: it sets CLAUDE_CODE_SIMPLE=1, and that same switch makes the CLI refuse to read OAuth (measured: duration_api_ms 0 and an authentication error). The two cannot be separated, so such a run got the whole prompt -- and the whole prompt is mostly schemas. Asked to list its tools, one of these sessions names forty-two of them: Artifact, CronCreate, DesignSync, PushNotification, Workflow, ShowOnboardingRolePicker and the rest, none reachable from a throwaway worktree, all priced every turn. --allowedTools does not help; it auto-approves, and every other schema is sent regardless. What decides which schemas exist is which tools are defined, and --agents defines its own. Measured on one endpoint with the same one-line prompt: 24 550 tokens of context per turn with --allowedTools, 6 522 with an --agents entry declaring five tools, 8 042 for the port's own executor command (three more tools, Bash among them), against 4 868 under --bare. Two thirds of the tax, gone, with the sign-in intact. lean_agent_flags returns nothing when --bare is already in play, so a keyed run still takes the better path and the declaration is only what a sign-in gets instead. The remaining gap to --bare is the system prompt, which --agents cannot touch. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
…andbox A key is a string in the environment and crosses into a container with it, which is why a keyed run is isolated and authenticated at once. A sign-in is not a string: the CLI reaches the endpoint through the host's session ingress, which the container does not have and which this module does not put there. So every containerised session answered "Not logged in · Please run /login", and the run found out one episode at a time -- sessions=4 failed=4 edits=0, architect designed 0 nodes, a whole run spent on a condition that was knowable before the first episode started. SessionSandbox now reports it beside the two reasons it already gave, a missing container engine and an unmountable toolchain, and degrades to a plain directory with the reason printed. The remedy it names is an API key, which is also the arrangement with the smallest prompt, since a key is what --bare needs. Running signed-in without the sandbox is the other option and is worth reading carefully: the session is then on the host as the host's user and reaches everything the sign-in reaches anyway, so what that trade gives up is the isolation, not the exposure. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
A rollout descends the entire Context Tree -- the root delegates, every leaf writes, the parents fold, and one proposal carries all of it. Two bounds stood between that and the accepted version. SpatialContract.max_files_per_diff is the one raised to 64 for a session executor; RecursiveDelegation.max_edits sits upstream of it in _bound and was 4, a trust region sized for a single completion proposing a file or two. Only the tighter one ever applied, so raising the other changed nothing. Measured on an 8-node fly tree: the sessions wrote 153 implementation files, every sweep committed exactly 4, and after three rollouts the accepted state held nine records and six Python files. src/training wrote 26 of them, src/brain/circuits 24, src/arena 22, src/frontend 21 -- and the version grew by four a round. Both now read one named value. What _bound does within the bound is unchanged and was not a defect: a node-creating record is trimmed last, because nothing else in the accepted version says the node exists and the source file it would be dropped for is re-proposable next round; work next; routine upkeep first. A test pins that order, and another pins the two caps to one name so they cannot drift apart again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
A proposal crosses five bounds between the session that writes a file and the state that keeps it. Two of them were this port's and were raised together. Two are the engine's -- `trust_region_ops` and `trust_region_chars` -- and nothing here had ever passed an `agg_config`, so the tightest cap on the path was a default of six ops, sized for a rule table and applied to a proposal carrying an eight-node repository. Measured: one rollout, thirteen sessions, 680 turns, forty minutes, merged `committed=0 rejected=1 [oversized=1]` with `discarded_diffs=0` and `truncated_edits=0` -- the port's own bounds passed it through cleanly and the engine rejected it behind them. Nine CONTEXT.md records landed and not one line of implementation. The largest single file any session wrote was 12,813 chars, so only the op count was ever binding. `engine_bounds(strategy)` derives the engine's two from the strategy that carries the cap, so there is one gate and it is the one that counts what it drops. `batch_trigger` and `max_wait_rounds` are carried across because `evolve()` builds `AggregatorConfig(batch_trigger=2, max_wait_rounds=1)` when passed nothing, not the dataclass defaults, and a config passed in replaces that object whole -- a run widening its trust region would otherwise have doubled its batch trigger in silence. Three tests, each verified to fail against the old behaviour: the caps all read one number, a diff at the full cap clears the engine's real trust-region predicate, and the batching numbers are read back out of a live `evolve()` rather than repeated as literals. `agentdescent/` is unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
Birfy
marked this pull request as draft
October 10, 2026 13:09
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Follow-up to #185, which merged.
agentdescent/is still untouched.#185 shipped the port and measured it on four library domains with oracle or frozen-suite scoring. This PR is what a fifth domain — a product, grown by a third-party model through real Claude Code sessions — exposed. Most of the commits are defects a run found, each one costing a wrong number before it was understood, and each commit message carries the run that produced it.
1. Blind mode: the agents write the tests
The port had upstream's test relationship backwards. Here a human wrote the assertions and the failing source was pasted into the executor's prompt. Upstream's agents write their own tests — "tests are the definition of done" — and c-testsuite / LLVM / Csmith are the experimenter's external benchmark, applied afterwards.
So the acceptance suite now lives outside the repository (
TestSuite.hidden) and the prompt on failure is the requirement text plus the assertion's name, never its source:Two things a frozen suite cannot say are now enforced by
own_review, inside the episode, by the parent:own_test_failuresruns those agent-written tests in one interpreter — themix testupstream's manager actually runs.2.
fly: a Drosophila brain, three assays, and a page to watch it learn onThe fifth domain, and the first with no reference implementation and no hand-written tree. The repository starts with one frozen file,
REQUIREMENTS.md, written in Chinese; everything else — theCONTEXT.mdtree, the code, the tests, andfly.pyitself — is grown.The requirement is connectome-grounded rather than a toy: AL divisive normalisation (Olsen–Wilson 2010), ~5% KC sparse coding via APL, MB compartments with the γ1pedc/PPL1 and γ5β′2a/PAM pairs, depression-only plasticity with timing-dependent sign (Handler 2019), MBON→DAN feedback (Felsenberg 2018), lateral-horn innate valence, and a CX ring attractor (EPG wedges, PEN shift, Δ7 inhibition, PFL3→DNa02 steering).
fly.pydeliberately is not frozen. #185's own deviation table called the frozenjqxentry point "a choice, and a defensible one" — not something upstream has. Upstream's contract is implicit in the objective: "build a C compiler" pinscc -o foo foo.c, which is why c-testsuite runs at all. "A fly simulator" pins nothing, so the calling convention has to be said — but it belongs to the requirement ("how I will invoke it"), not to the repository ("here is code for you"). It is now §4 ofREQUIREMENTS.md, and the agents write the file.3. Over-decomposition, and the threshold that fixed it
An architect given a product built 600 nodes at depth 8, 75% of them routers that forward and do nothing —
apl, one neuron, was given two children.The fix is a declaration, not a cap: the architect must state how many files each child holds, and a child of two files or fewer is not a directory. Trees then came out at 14–25 nodes and depth 3–5 — upstream's archive shows 26 nodes and observed depth 5 — and converged on their own, never reaching the node budget.
Three more Phase 1 defects, all found by watching a tree come out wrong:
…/cell_types/mushroom_bodyredesigned the whole library from the top_apply, and resume was decided after dispatch — spending the call it exists to savePhase 1 now designs a level at a time over a state snapshot, and prints each node as it lands.
4. An executor session that can reach a third-party endpoint
--executor claude-codespawned the CLI with an inherited environment. Inside a managed Claude Code session that environment carriesCLAUDE_CODE_REMOTE, which puts the CLI on the host's session ingress and makes it ignore theANTHROPIC_BASE_URLand key it was handed. Every episode sent the host's token to a third-party endpoint and got 401: one run opened 28 sessions, failed 23, and wrote no files at all.What the CLI reports for this is
Authentication error · This may be a temporary network issue, please try again— neither — and the stderr beside it names an unrecognised model forquery_source: generate_session_title, which is the session-title side query, not the agent turn, and prints identically on runs that work. I read that line as the cause and was wrong; bisection over the environment showedCLAUDE_CODE_REMOTEis the single switch.cli_envdrops those variables, but only when the run actually points somewhere else; a session on the host's own provider is untouched. The credential only ever travels from the environment into a child'senv=.Two smaller ones beside it: an executor session now runs the model the run asked for (it previously took the CLI's default unless
--providerhappened to beclaude-cli, so a run could be on two models in silence), and--max-turnsno longer overrides upstream's 2048-root / 128-child split with its own parser default.5. Two ways a run reported success it had not earned
An empty suite is not a passing suite.
own_test_failuresreturns[]when the repository holds no test files, which reads as "everything passes" — and that answer is the precondition forcomplete_task. A root agent used it: sixteen implementation files, zero tests, acceptance reward 0.000, and it claimed the objective delivered five times and was believed every time. That is backwards from the rule the same class already enforces on every child;require_testsapplies it wherever the answer is a gate, and stays off where the question is really "did anything break".A reasoning model's thinking is spent from
--max-tokenstoo. The architect asked for a repository design and got 1708 characters of JSON that stopped mid-word. The run reportedarchitect designed 0 nodes, 1 replies unusableand then grew an entire phase 2 with no tree at all — a shape that looks like a bad model or a parser bug and is neither. Measured on the architect prompt: 4096 emits ~15k output tokens across attempts in 148s and truncates about half the time; 16000 takes 264s and parses. Disabling thinking is the other lever and the worse one — it changes what the model produces, not only how much.--max-tokensand--timeoutare now declared and honoured inexamples/_common.py, which is what that module exists for. They arrive together because raising one alone only moves the failure: the call that now fits takes 264s, andclaude()'s 120s default would have timed it out three times over. Both are opt-in by naming a default — nine ports already declare--max-tokensthemselves with numbers they measured (4096, 16000, 32000), so declaring it unconditionally is an argparse conflict that takes those entry points down at import. That is how I found them: six tests acrossporous,openevolveandmetasearch, caught before this was pushed.6. Session isolation: a session launched from a session is not that session
A session spawned from inside a Claude Code session inherits that session's equipment, and equipment is priced per turn whether or not it is reachable. Measured against one endpoint with the same one-line prompt: 31 850 input tokens plain, 31 099 with this port's own permission flags, 1 317 bare. The middle number is the mechanism — permission flags say what may be called, and every other schema is sent anyway. From the transcript's own snapshot the system prompt is 5 720 characters and the tool schemas are 178 742, of which
Artifactalone is 64 168.Identity leaked with it. Inherited, every session in a run is the host: one run's episodes each wrote a transcript named with the host's session id, and their
TodoWritestate — keyed by that id — landed in the host's task list, two hundred entries of "Implement src/brain package".HOST_SESSION_VARSdrops the identity variables andCLAUDE_CONFIG_DIRmoves the whole CLI state directory out of~/.claude.--bareis what removes the schemas, and it has a trap: it also removes theWritetool, and--allowedToolscannot put it back (that flag is an auto-approve list, not a whitelist — a manager reached forEdit41 times without it ever being listed).available_toolsswapsWriteforEdit, which creates files that do not exist.The sandbox is the engine's own
sandbox_container: docker/podman, read-only root,--cap-drop ALL, no network unless asked, one CLI home per container. It exists because the prompt is a rule and not a wall — four of twelve episodes in one run ranfind /and read a previous run's output from/tmp, one of them opening the very_cli.pythat answered the acceptance failure it had been asked to reproduce.7. Four the throughput of a real run exposed
The file-count cap was eating whole episodes.
SpatialContractcaps a proposal atmax_files_per_diff, and the cap rejects the diff, not the files over the line — so an episode one file past it contributes nothing. At 6, counting distinct paths written across 309 productive sessions, 11% of the episodes that did work lost all of it; the largest was carrying 23 files. Nothing said so: the two existing counters are about authority (wrote outside its subtree, mistook a file for a node) and neither moves for this, so a run that lost an eighth of its work and a run whose agents had nothing to say printed the same header. The drop is now counted (discarded_diffs=N (M files)), and the cap stops pretending to be a trust region — the trust region is the node's subtree, enforced edit by edit, and twelve files under one node are not more dangerous than six. For a session executor it is a runaway guard at 64; a single completion asked for whole files keeps 6, because there six really is one.The blind failure read as an order at every node. It ended "write your own test that reproduces this, put it beside the code it covers, and make both pass". Every agent in a rollout is shown the same text, so at a leaf it reads as an instruction to fix the import here: eight different nodes each wrote their own
_cli.py, 79 writes between them,src/brain/olfactory/_cli.pyalone 21 times. Only the one at the repository root could ever have resolved it. The evidence now says it is the repository's, not necessarily yours — judge it against what you own, and name the path in the final message if the fix belongs elsewhere.The growth phase printed nothing. Between
Growing the worldand the final summary the run was silent, and for a formation run that is hours. Finding out whether anything was being accepted meant reading CLI transcripts and looking inside live containers — and a live worktree holds the base state plus whatever the session has written since, which is not the accepted state, so readings drawn that way are wrong more often than not.evolve()already takes anon_roundhook; the run now passes one and prints a line per merger sweep.RoundInfo.reasonsgoes on the end becausecommitted=0has two causes that need opposite fixes: the gate refused the work, or the work never reached the gate.A rollout was an hour, and the workers spent it on the same tree.
RecursiveDelegation.proposeis a walk of the whole Context Tree with one Claude Code session per node, and it was aforloop. On an 8-node tree at five to fifteen minutes a session that is about an hour per rollout —py-spyon a live run showed all four workers still inside their firstpropose()after 75 minutes, three levels deep in nested_episodeframes, with three files in the accepted state. That reads as an acceptance problem and is not one: the parent gate accepts a tie, four of the five reviews that had run said ACCEPT, and the staleness policy keeps a card whose reward is merely unchanged. Nothing was being rejected; nothing had finished. And--workerscounts whole tree walks, so four workers were four independent descents of the same eight nodes — two of them onsrc/brain/navigationat the same moment, doing the same node's work, of which one result could survive.Siblings now run concurrently and are folded in
delegationsorder, so the proposal a rollout returns is unchanged. The spatial contract is what makes that safe rather than a merge problem: a child writes only under its own subtree, so two siblings cannot touch the same path. The knobs say which parallelism you are asking for:--workers--node-workers--max-sessionsThe bound lives in
run_clirather than in the thread pools: sessions are the scarce thing, a pool per level would multiply, and sessions never nest — a manager's own session runs before and after its children's, never during — so one semaphore there covers the whole tree with no way for a parent to deadlock on its own children. Every counter the port reports is now incremented under a lock;x += 1is a read and a write with a bytecode boundary between them, and these numbers are the port's evidence.8. What is not claimed
Runs died on endpoint problems rather than code, and every one of them first looked like a code bug:
CLAUDE_CODE_REMOTE(§4)10 nodes, 9 replies unusableexcept Exceptioncounting it as a parse failureAccountQuotaExceededagainst 53 ×AccountRateLimitExceededSo there is no end-to-end
flynumber in this PR, and none is claimed. The domain, the blind mode and every fix above are exercised by the offline tests; the product itself has not been grown to completion on a healthy endpoint.What the last run did confirm before the quota ran out: the sibling parallelism (concurrency 1 → 6, six sessions on six distinct nodes, no two on the same node, the semaphore holding at its bound) and the isolation fix (4 868 input tokens on a first turn against 31 850 before it,
session_contextempty, no skill listing, agent listing or user identity in the prompt).One question is deliberately left open: an earlier run reported
claude code edits=298but landed 37 files withtruncated_edits=56, and an endpoint dying mid-run produces the same shape, so whether that gap is a real defect in the edit-collection path can only be judged on a run that stays healthy.The
md,minilangandstackvmnumbers in #185 are unchanged by this PR.Verification
pytest -q # full suite green python -m examples.genesis.genesis_recursive_worlds --domain fly --dry-run python -m examples.genesis.genesis_recursive_worlds --domain fly --architect --complete-task \ --executor claude-code --workers 1 --node-workers 6 --max-tokens 16000 --timeout 600 --model ...tests/test_genesis_example.pycovers blind mode,own_review, the file-count threshold, both routing-table directions, resume,cli_env, the empty-suite gate, the diff-cap discard and its counter, the growth-phase progress hook, sibling concurrency with an ordered fold, and the session bound.--dry-runcrosses no boundary and the default run needs no key.🤖 Generated with Claude Code
https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW