Skip to content

Genesis: blind mode, a product domain, and eighteen defects a real run found - #188

Draft
Birfy wants to merge 38 commits into
mainfrom
claude/agentdescent-evoxgenesis-fw6wix
Draft

Birfy wants to merge 38 commits into
mainfrom
claude/agentdescent-evoxgenesis-fw6wix

Conversation

@Birfy

@Birfy Birfy commented Sep 15, 2026 •

Copy link
Copy Markdown
Owner

Follow-up to #185, which merged. agentdescent/ is still untouched.

#185 shipped the port and measured it on four library domains with oracle or frozen-suite scoring. This PR is what a fifth domain — a product, grown by a third-party model through real Claude Code sessions — exposed. Most of the commits are defects a run found, each one costing a wrong number before it was understood, and each commit message carries the run that produced it.

1. Blind mode: the agents write the tests

The port had upstream's test relationship backwards. Here a human wrote the assertions and the failing source was pasted into the executor's prompt. Upstream's agents write their own tests — "tests are the definition of done" — and c-testsuite / LLVM / Csmith are the experimenter's external benchmark, applied afterwards.

So the acceptance suite now lives outside the repository (TestSuite.hidden) and the prompt on failure is the requirement text plus the assertion's name, never its source:

{requirement heading}

ACCEPTANCE: the trained animal turns toward the rewarded odour.

Two things a frozen suite cannot say are now enforced by own_review, inside the episode, by the parent:

  • a child whose own tests fail has not finished, whatever the hidden suite thinks;
  • a child that wrote an implementation file and no test anywhere under its node has not finished either.

own_test_failures runs those agent-written tests in one interpreter — the mix test upstream's manager actually runs.

2. fly: a Drosophila brain, three assays, and a page to watch it learn on

The fifth domain, and the first with no reference implementation and no hand-written tree. The repository starts with one frozen file, REQUIREMENTS.md, written in Chinese; everything else — the CONTEXT.md tree, the code, the tests, and fly.py itself — is grown.

The requirement is connectome-grounded rather than a toy: AL divisive normalisation (Olsen–Wilson 2010), ~5% KC sparse coding via APL, MB compartments with the γ1pedc/PPL1 and γ5β′2a/PAM pairs, depression-only plasticity with timing-dependent sign (Handler 2019), MBON→DAN feedback (Felsenberg 2018), lateral-horn innate valence, and a CX ring attractor (EPG wedges, PEN shift, Δ7 inhibition, PFL3→DNa02 steering).

fly.py deliberately is not frozen. #185's own deviation table called the frozen jqx entry point "a choice, and a defensible one" — not something upstream has. Upstream's contract is implicit in the objective: "build a C compiler" pins cc -o foo foo.c, which is why c-testsuite runs at all. "A fly simulator" pins nothing, so the calling convention has to be said — but it belongs to the requirement ("how I will invoke it"), not to the repository ("here is code for you"). It is now §4 of REQUIREMENTS.md, and the agents write the file.

3. Over-decomposition, and the threshold that fixed it

An architect given a product built 600 nodes at depth 8, 75% of them routers that forward and do nothing — apl, one neuron, was given two children.

The fix is a declaration, not a cap: the architect must state how many files each child holds, and a child of two files or fewer is not a directory. Trees then came out at 14–25 nodes and depth 3–5 — upstream's archive shows 26 nodes and observed depth 5 — and converged on their own, never reaching the node budget.

Three more Phase 1 defects, all found by watching a tree come out wrong:

symptom cause
five brain regions silently vanished on resume the routing table only dropped refused children, never added forgotten ones — and one node's table held API-Surface prose
an architect at …/cell_types/mushroom_body redesigned the whole library from the top resume handed every child the parent's objective instead of its own routing line
every node cost a model call twice parallelising the level re-asked inside _apply, and resume was decided after dispatch — spending the call it exists to save

Phase 1 now designs a level at a time over a state snapshot, and prints each node as it lands.

4. An executor session that can reach a third-party endpoint

--executor claude-code spawned the CLI with an inherited environment. Inside a managed Claude Code session that environment carries CLAUDE_CODE_REMOTE, which puts the CLI on the host's session ingress and makes it ignore the ANTHROPIC_BASE_URL and key it was handed. Every episode sent the host's token to a third-party endpoint and got 401: one run opened 28 sessions, failed 23, and wrote no files at all.

What the CLI reports for this is Authentication error · This may be a temporary network issue, please try again — neither — and the stderr beside it names an unrecognised model for query_source: generate_session_title, which is the session-title side query, not the agent turn, and prints identically on runs that work. I read that line as the cause and was wrong; bisection over the environment showed CLAUDE_CODE_REMOTE is the single switch.

cli_env drops those variables, but only when the run actually points somewhere else; a session on the host's own provider is untouched. The credential only ever travels from the environment into a child's env=.

Two smaller ones beside it: an executor session now runs the model the run asked for (it previously took the CLI's default unless --provider happened to be claude-cli, so a run could be on two models in silence), and --max-turns no longer overrides upstream's 2048-root / 128-child split with its own parser default.

5. Two ways a run reported success it had not earned

An empty suite is not a passing suite. own_test_failures returns [] when the repository holds no test files, which reads as "everything passes" — and that answer is the precondition for complete_task. A root agent used it: sixteen implementation files, zero tests, acceptance reward 0.000, and it claimed the objective delivered five times and was believed every time. That is backwards from the rule the same class already enforces on every child; require_tests applies it wherever the answer is a gate, and stays off where the question is really "did anything break".

A reasoning model's thinking is spent from --max-tokens too. The architect asked for a repository design and got 1708 characters of JSON that stopped mid-word. The run reported architect designed 0 nodes, 1 replies unusable and then grew an entire phase 2 with no tree at all — a shape that looks like a bad model or a parser bug and is neither. Measured on the architect prompt: 4096 emits ~15k output tokens across attempts in 148s and truncates about half the time; 16000 takes 264s and parses. Disabling thinking is the other lever and the worse one — it changes what the model produces, not only how much.

--max-tokens and --timeout are now declared and honoured in examples/_common.py, which is what that module exists for. They arrive together because raising one alone only moves the failure: the call that now fits takes 264s, and claude()'s 120s default would have timed it out three times over. Both are opt-in by naming a default — nine ports already declare --max-tokens themselves with numbers they measured (4096, 16000, 32000), so declaring it unconditionally is an argparse conflict that takes those entry points down at import. That is how I found them: six tests across porous, openevolve and metasearch, caught before this was pushed.

6. Session isolation: a session launched from a session is not that session

A session spawned from inside a Claude Code session inherits that session's equipment, and equipment is priced per turn whether or not it is reachable. Measured against one endpoint with the same one-line prompt: 31 850 input tokens plain, 31 099 with this port's own permission flags, 1 317 bare. The middle number is the mechanism — permission flags say what may be called, and every other schema is sent anyway. From the transcript's own snapshot the system prompt is 5 720 characters and the tool schemas are 178 742, of which Artifact alone is 64 168.

Identity leaked with it. Inherited, every session in a run is the host: one run's episodes each wrote a transcript named with the host's session id, and their TodoWrite state — keyed by that id — landed in the host's task list, two hundred entries of "Implement src/brain package". HOST_SESSION_VARS drops the identity variables and CLAUDE_CONFIG_DIR moves the whole CLI state directory out of ~/.claude.

--bare is what removes the schemas, and it has a trap: it also removes the Write tool, and --allowedTools cannot put it back (that flag is an auto-approve list, not a whitelist — a manager reached for Edit 41 times without it ever being listed). available_tools swaps Write for Edit, which creates files that do not exist.

The sandbox is the engine's own sandbox_container: docker/podman, read-only root, --cap-drop ALL, no network unless asked, one CLI home per container. It exists because the prompt is a rule and not a wall — four of twelve episodes in one run ran find / and read a previous run's output from /tmp, one of them opening the very _cli.py that answered the acceptance failure it had been asked to reproduce.

7. Four the throughput of a real run exposed

The file-count cap was eating whole episodes. SpatialContract caps a proposal at max_files_per_diff, and the cap rejects the diff, not the files over the line — so an episode one file past it contributes nothing. At 6, counting distinct paths written across 309 productive sessions, 11% of the episodes that did work lost all of it; the largest was carrying 23 files. Nothing said so: the two existing counters are about authority (wrote outside its subtree, mistook a file for a node) and neither moves for this, so a run that lost an eighth of its work and a run whose agents had nothing to say printed the same header. The drop is now counted (discarded_diffs=N (M files)), and the cap stops pretending to be a trust region — the trust region is the node's subtree, enforced edit by edit, and twelve files under one node are not more dangerous than six. For a session executor it is a runaway guard at 64; a single completion asked for whole files keeps 6, because there six really is one.

The blind failure read as an order at every node. It ended "write your own test that reproduces this, put it beside the code it covers, and make both pass". Every agent in a rollout is shown the same text, so at a leaf it reads as an instruction to fix the import here: eight different nodes each wrote their own _cli.py, 79 writes between them, src/brain/olfactory/_cli.py alone 21 times. Only the one at the repository root could ever have resolved it. The evidence now says it is the repository's, not necessarily yours — judge it against what you own, and name the path in the final message if the fix belongs elsewhere.

The growth phase printed nothing. Between Growing the world and the final summary the run was silent, and for a formation run that is hours. Finding out whether anything was being accepted meant reading CLI transcripts and looking inside live containers — and a live worktree holds the base state plus whatever the session has written since, which is not the accepted state, so readings drawn that way are wrong more often than not. evolve() already takes an on_round hook; the run now passes one and prints a line per merger sweep. RoundInfo.reasons goes on the end because committed=0 has two causes that need opposite fixes: the gate refused the work, or the work never reached the gate.

A rollout was an hour, and the workers spent it on the same tree. RecursiveDelegation.propose is a walk of the whole Context Tree with one Claude Code session per node, and it was a for loop. On an 8-node tree at five to fifteen minutes a session that is about an hour per rollout — py-spy on a live run showed all four workers still inside their first propose() after 75 minutes, three levels deep in nested _episode frames, with three files in the accepted state. That reads as an acceptance problem and is not one: the parent gate accepts a tie, four of the five reviews that had run said ACCEPT, and the staleness policy keeps a card whose reward is merely unchanged. Nothing was being rejected; nothing had finished. And --workers counts whole tree walks, so four workers were four independent descents of the same eight nodes — two of them on src/brain/navigation at the same moment, doing the same node's work, of which one result could survive.

Siblings now run concurrently and are folded in delegations order, so the proposal a rollout returns is unchanged. The spatial contract is what makes that safe rather than a merge problem: a child writes only under its own subtree, so two siblings cannot touch the same path. The knobs say which parallelism you are asking for:

flag what it multiplies
--workers whole tree walks at once (keep small; they duplicate each other)
--node-workers siblings one manager runs at once
--max-sessions hard bound on what the endpoint sees; defaults to their product

The bound lives in run_cli rather than in the thread pools: sessions are the scarce thing, a pool per level would multiply, and sessions never nest — a manager's own session runs before and after its children's, never during — so one semaphore there covers the whole tree with no way for a parent to deadlock on its own children. Every counter the port reports is now incremented under a lock; x += 1 is a read and a write with a bytecode boundary between them, and these numbers are the port's evidence.

8. What is not claimed

Runs died on endpoint problems rather than code, and every one of them first looked like a code bug:

symptom actual cause
23/28 sessions failed, 0 edits CLAUDE_CODE_REMOTE (§4)
258/276 sessions failed, 1293/2009 model calls failed 402 Insufficient Balance — the account ran out mid-run
10 nodes, 9 replies unusable 429 AccountQuotaExceeded — a 5-hour quota, except Exception counting it as a parse failure
the last run stopped after phase 2 began 78 × AccountQuotaExceeded against 53 × AccountRateLimitExceeded

So there is no end-to-end fly number in this PR, and none is claimed. The domain, the blind mode and every fix above are exercised by the offline tests; the product itself has not been grown to completion on a healthy endpoint.

What the last run did confirm before the quota ran out: the sibling parallelism (concurrency 1 → 6, six sessions on six distinct nodes, no two on the same node, the semaphore holding at its bound) and the isolation fix (4 868 input tokens on a first turn against 31 850 before it, session_context empty, no skill listing, agent listing or user identity in the prompt).

One question is deliberately left open: an earlier run reported claude code edits=298 but landed 37 files with truncated_edits=56, and an endpoint dying mid-run produces the same shape, so whether that gap is a real defect in the edit-collection path can only be judged on a run that stays healthy.

The md, minilang and stackvm numbers in #185 are unchanged by this PR.

Verification

pytest -q                                            # full suite green
python -m examples.genesis.genesis_recursive_worlds --domain fly --dry-run
python -m examples.genesis.genesis_recursive_worlds --domain fly --architect --complete-task \
    --executor claude-code --workers 1 --node-workers 6 --max-tokens 16000 --timeout 600 --model ...

tests/test_genesis_example.py covers blind mode, own_review, the file-count threshold, both routing-table directions, resume, cli_env, the empty-suite gate, the diff-cap discard and its counter, the growth-phase progress hook, sibling concurrency with an ordered fold, and the session bound. --dry-run crosses no boundary and the default run needs no key.

🤖 Generated with Claude Code

https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW

`fly`: a Drosophila brain, three assays, an HTTP backend and the page a
person watches it learn on. The four domains before it grow a library; this
one is the first asked for a frontend, so the specification alone has to carry
the design.

It is also the first shipped with no reference implementation, and the cost is
stated rather than hidden: no offline actor, so `--domain fly` without a model
now says so instead of raising AttributeError, and no kill_report, so nothing
in the repository checks what this suite rejects. What stands in its place is
that every number in the specification was measured on a throwaway
implementation while the suite was being written -- 91 assertions passing, and
three of the specification's least obvious paragraphs are there because that
implementation failed them first: a ring attractor sharpened with any pointwise
nonlinearity quantises heading to the nearest wedge; an avoidance assay with
the punished source in the middle of a walled arena measures the arena, not the
animal; and `src/learn/` beside a public `learn()` is two things called
`src.learn`, where the submodule import rebinds the name.

That last one is the same collision this port was bitten by twice from the
other direction, which is why the specification names it and
test_integration.py asserts it.

93 tasks over 8 test files, 10 of them an audit set no agent can read.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
…e suite

Upstream the agents write the tests. They have file and shell tools, they
write `test/*.exs`, and "tests are the definition of done"
(agents/manager.ex); the external suites Genesis validates against --
c-testsuite, LLVM, Csmith -- are the experimenter's measurement, applied
afterwards. This port had it backwards. A human wrote every assertion, and
then `TEST_FAILURE` pasted its **source** into the prompt, which is the
strongest hint there is: an agent handed the assertion is not implementing a
specification, it is writing to an assertion.

`TestSuite.hidden` is the other way round. The acceptance suite never enters
the repository -- it is merged into the scratch copy one evaluation sees, the
way the audit set already was -- and the prompt becomes the requirement it was
written from plus the assertion's name. The name carries real information and
that is deliberate: `test_the_code_is_sparse_at_every_concentration` is a
requirement written as a sentence. The threshold, the tolerance and the inputs
are exactly what it does not carry.

So every test *in* the repository is an agent's own, and two new things can be
asked of a child. `own_test_failures` runs them in one interpreter -- `mix
test`, theirs -- scoped to one node's subtree. `own_review` refuses a child
whose own tests fail, and a child that wrote an implementation file and no
test for it: upstream every node is accountable for its own subtree, so a node
with untested code is a node whose parent has nothing to run.

Also `ConcurrencyGauge`, because the obvious arithmetic is wrong.
`usage.seconds / wallclock` looks like it says how many calls were in flight
and does not -- `seconds` spans phases that run before `evolve()` does, while
a stage profile's `wallclock` covers only the stage, and the md run's ratio
came out at 8.2 with four workers. It counts them directly instead, and the
run reports the peak. In `examples/`, not in the engine: what the engine's
concurrency *is* is its own business; what an example observed is the
example's.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
…he rest

The domain now hands over two things and both are requirements rather than
design. `REQUIREMENTS.md` is what the user asked for, in their words -- a
connectome-grounded Drosophila brain, three assays, learning that visibly
changes behaviour, a frontend and a backend, and a test beside every
directory's code -- plus a reading list. No equations, no coefficients, no
module layout. `fly.py` is that requirement made executable, and it imports
exactly four names from `src`, so the public surface is four functions and
everything else is the agents' internal business.

Everything else grows: the CONTEXT.md tree, the circuit, the layout, the
frontend, and every test in the repository.

Both assertion sets stay outside it. `hidden` is black-box acceptance and
drives the search -- the agent is shown the requirement heading, the
assertion's name as a sentence, and what it reported, never a line of source.
`audit` is sealed: written before the run, never in the repository, the same
requirements asked with different seeds and different tasks, and run once at
the end. The gap between "acceptance passes" and "the sealed suite passes" is
the measurement of whether it built the thing or wrote to the acceptance.

`suite_failures` now asks what upstream's manager asks -- do *your own* tests
pass -- because the human suite is not something this domain's agents can see.
`own_review` joins the review chain, so a child that wrote an implementation
file and no test for it is sent back.

One requirement neither set can check, stated rather than hidden: whether the
brain is really built from the connectome. Black-box cannot ask. Only the
parent reading the diff, and a person reading the result, can.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
…inary

The SDK path needs an `ANTHROPIC_API_KEY`. The CLI is authenticated another
way entirely -- an OAuth session, a subscription, a corporate login -- and on
a machine where that is the credential that exists there was no way to run a
port at all. This session hit exactly that: the endpoint returned `401
ModelArts.81015, Invalid client ip ... rejected by api key ip whitelist
setting`, and `preflight` caught it in one call rather than in 400 episodes of
silent nothing.

`--provider claude-cli` is a one-shot `claude -p` with every tool denied,
which is a plain prompt-to-text function and so is exactly what a Completion
is. Two caveats are in the docstring rather than hidden, because both change
the output and not only the plumbing: the CLI wraps the prompt in its own
system prompt, tool definitions and any CLAUDE.md it finds, so a call carries
tens of thousands of cached input tokens the API path would not; and the cost
lands on the CLI's credentials, where no budget flag here can see it. Token
accounting counts cache reads and writes as the prompt tokens they are --
`input_tokens` alone reported 10 for a call that really carried 28,000. And
`--no-thinking` says it has no effect here instead of being dropped in
silence.

Two fixes the same run needed. The Claude Code executor now runs the model the
run asked for, so the architect and every executor episode are not on
different models by default. And its failure template comes from the domain
rather than a constant: handing a blind domain's session the assertion's
source is the one thing that domain exists to prevent.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
… budget

Three of the four gaps the deviations table recorded, closed.

**Upstream's architect Phase 3.** It does not stop at design: it reviews the
implementation and re-spawns refinement architects where a node misaligns
(`agents/architect.ex`). This port stopped after the design, so a record
written before any code existed stayed the map for ever. `--refine` checks a
node's record against what is actually in the directory -- a routing table
promising a child nobody created, an `## API Surface` that never mentions a
file sitting right there -- and where it has drifted, re-spawns an architect
on that one node to rewrite it against the code as it is.

**And that is also the 62 updates.** Upstream's archive shows 26 `CONTEXT.md`
creations and 62 later accepted updates affecting 19 files: a record is
maintained, not written once. This port wrote them on two occasions, and a
`--mode a` run sat at 0.938 for 30 003 rollouts reading a map of a layout the
work had already left behind, with nothing in the mechanism able to say so.
The hook fires on the leaf path as well as the accountability pass, and leaves
need it more -- a leaf is where the code lands, so its API Surface is the
first thing to go stale.

**One episode, two budgets.** Upstream gives a root agent up to 2,048
model-tool turns and a child 128. The Claude Code executor had one number, 24,
for both, which either starves the root or hands every leaf a session it has
no use for. It now carries upstream's own two numbers, and a single
`--max-turns` still means a single number for a cheap run.

**And the node budget.** The fly domain's architect named 17 children it never
reached at the default of 12, so the tree came out two deep and truncated.
`--nodes` sets it.

The fourth gap, the human merge on the dashboard, stays recorded and unbuilt.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
…ing it

The executor learned upstream's two budgets -- 2048 at the root, 128 below --
and then the parser's own default of 24 fed straight past them, so a run
configured for the split reported "up to 24 turns" and meant it. The default
is now 0, which means "use the split", and a number given explicitly still
pins both to it for a cheap run. The header prints both numbers rather than
one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
The fly domain's architect spent 27 minutes designing 32 nodes and was about
to hit `--nodes 40` with a queue still behind it. Raising the budget meant
paying for all 32 again, because `design()` asked for every node it reached
whether or not a record was already there.

With `resume`, a node that already carries a record is kept: its routing table
is taken as the design, its children are queued, and no call is spent. The
thing that revises a record once code exists is the refinement architect, not
this.

Off by default, and deliberately. At `root_path=""` the record already in
`given` is the *harness* record -- generated, not designed -- so reusing it
would skip the root and leave the tree with no design at all. The runner turns
it on only for `--continue-from`, where a record that is there really is an
earlier phase 1's work.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
The routing line's right-hand side is not decoration -- it is the objective
the parent's architect handed that child. A phase 1 continuing from records
rather than from replies has nothing else to give one, and the first version
passed down the *parent's* objective instead.

What that does is visible in the run that found it. An architect asked to
design `src/brain/circuits/cell_types/mushroom_body`, while carrying the root
objective "implement the software REQUIREMENTS.md asks for", came back with
children `brain, learning, environment, simulation` -- it had redesigned the
whole library from the top, at depth five, three nodes running.

`parse_routes` returns `(path, what it handles)` where `parse_routing` returns
the paths alone, and resume queues each child with its own line.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
…ctions

Upstream calls the routing table "your primary delegation tool ... the map
that makes recursive delegation work". A record whose table disagrees with the
children the architect just opened is a broken map, and phase 1 now repairs it
rather than passing it on.

Dropping a refused entry was already there -- a table advertising a node
nobody may write is a trap for the next manager, which would delegate there
and be refused in turn. Adding a forgotten one turned out to matter more. An
architect answered with five children and a `## Routing Table` section holding
its API Surface instead: "`__init__.py` — Exports get_regions(...)", then two
sentences about FlyWire. Nothing looked wrong that run, because the children
were queued from the *reply*. Then the budget ran out, a later phase 1 resumed
from records, rebuilt its queue from the **table**, found no routes in it, and
dropped an entire subtree of the brain -- five regions, silently, because the
record and the reply had disagreed and only the reply was ever right.

And a second guard from the same run. "If the objective feels too large, that
is exactly the signal to decompose MORE aggressively" has a counterweight
upstream states in the same breath: single responsibility, and shared
capability belongs at the lowest common ancestor. A node named after one of
its own ancestors is that rule broken in the one way a tree can show -- what
is really in there belongs to the ancestor, or the ancestor's name was wrong,
and either way two places claim it and no later agent can tell which. The run
produced `.../receptor/types/catalog/types` and `.../catalog/demographics`
beside an existing `.../receptor/population/demographics`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
… at a time

**The driver was a choice, and this port's own deviations table said so**: "a
frozen entry point, so the result is a program rather than a package nobody
can invoke. A choice, and a defensible one." Upstream Mode B takes an
objective and nothing else; c-testsuite, LLVM and Csmith are the
experimenter's benchmark, applied afterwards, never in the repository.

Upstream does have a contract -- it is just implicit in the objective. "Build
a C compiler" pins `cc -o foo foo.c`, an convention everyone already knows,
which is why c-testsuite can run at all. "A fruit fly simulator" has no such
convention, so it has to be said -- but it belongs in the *requirement* (how I
want to use it) rather than in the repository (here is code you may not
touch). `REQUIREMENTS.md` §4 now says which three subcommands must work and
that each takes `--format json`, because a person who wants to plot a learning
curve asks for machine-readable output. `fly.py` is theirs to write.

So both suites are black box now. They start `python fly.py ...`, read its
JSON, drive its HTTP server over a socket, and import nothing -- which is the
shape of upstream's own validation. One file is frozen: the requirement.

**And phase 1 designs a level at a time.** Upstream *spawns* sub-architects,
which is a statement about independence: every node at a depth inherits the
chain down to its own parent, designed a level ago, so nothing in a level can
depend on anything else in it. Running them one after another was this port's
choice and it cost the fly domain 32 minutes for 71 nodes. The asks read a
snapshot rather than the live tree, so a sibling can never see another
sibling's record even if it finishes first, and the tree does not depend on
which call returns when.

Two bugs the change surfaced, both fixed: `_apply` still carried the serial
loop's own `self._ask`, so every node was asked twice; and resume was decided
after the dispatch, which spends exactly the call it exists to save.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
A Claude Code episode goes through the local CLI whatever `--provider` says,
and the CLI was handed no model at all unless the provider happened to be
`claude-cli` -- so a run whose architect was on an API model had every
executor session on the CLI's default instead, two models, in silence.

It now takes the run's model, and `--executor-model` is there for the case
where they should deliberately differ. The CLI reads ANTHROPIC_BASE_URL and
ANTHROPIC_API_KEY from the environment, so pointing the SDK and the CLI at one
endpoint is all it takes to put both halves of a run on the same model.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
It printed nothing until the whole phase finished. On the fly domain that was
82 minutes of silence, and answering "how far in is it" meant parsing the
Claude CLI's own session log out of ~/.claude/projects. Then the architect
moved to the SDK, the subprocess and its log went away, and the question had
no answer at all -- six sockets open and a thread count.

Each node now prints as it lands, with the running count against the budget.
`--quiet-phase1` turns it off.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
Upstream says both things in the same breath -- decompose MORE aggressively
when the objective feels large, *and* single responsibility, shared capability
at the lowest common ancestor, a directory holding one short function is a
directory that did not want splitting. The first is a command. The second
needs judgement, and a model handed both executes the first.

Measured: the fly domain's architect designed 600 nodes for one simulator --
921 before the budget cut it -- three quarters of them pure routing, at depth
8. It gave `apl`, which is a single GABAergic neuron, two children and one of
those three more. Upstream's 123-hour C compiler run, 750 files and 249 000
physical lines, has **26** nodes and bottomed out at depth 5.

So the counterweight is enforced rather than stated. Each child now declares
`files` -- how many source files the architect expects that directory to hold,
counting everything below it -- and a child of two or fewer is refused: it is
two files in this node's API Surface, not a directory. A node of three or
fewer is told to return no children at all. An undeclared count is not a
refusal; the guard reads what is there rather than inventing a number.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
`_parse` now returns `files` on every child, undeclared as 0, and the
assertion that pinned the child dict verbatim had to say so.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
`--executor claude-code` spawned the CLI with an inherited environment. Inside
a managed Claude Code session that environment carries CLAUDE_CODE_REMOTE,
which puts the CLI on the host's session ingress and makes it ignore the
ANTHROPIC_BASE_URL and key it was handed. So every episode sent the host's
token to a third-party endpoint and got 401: one fly run opened 28 sessions,
failed 23, and wrote no files at all.

What the CLI reports for this is "Authentication error - This may be a
temporary network issue, please try again", which is neither, and the stderr
beside it names an unrecognized model for query_source generate_session_title
-- the session-title side query, not the agent turn, and printed just the same
on runs that work. I read that line as the cause and was wrong.

`cli_env` drops the host-provider variables, but only when the run actually
points somewhere else; a session on the host's own provider is untouched. The
credential still only ever travels from the environment into a child's env.

This also corrects 611c4af, which claimed pointing the SDK and the CLI at one
endpoint was all it took to put both halves of a run on the same model. It is
what the CLI documents, and it is true on a laptop. It was not true here, and
the port's own fly runs are what proved otherwise.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
`own_test_failures` returns [] when the repository holds no test files, which
reads as "everything passes" -- and that answer is the precondition for
`complete_task`. A fly root agent used it: sixteen implementation files, zero
tests, acceptance reward 0.000, and it claimed the objective delivered five
times and was believed every time.

That is backwards from the rule the same class already enforces on every child
in `own_review`: code with no test is not finished. `require_tests` applies it
to the whole repository, and the blind domain's `suite_failures` passes it,
because that path is a gate. It stays off where the question really is "did
anything the agents wrote break" -- a parent reviewing a child that has
legitimately returned no code yet should not be told its tests failed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
The architect asked for a repository design and got 1708 characters of a JSON
object that stopped mid-word. The run reported "architect designed 0 nodes,
1 replies unusable" and then grew a whole phase 2 with no tree at all, which
is a shape that looks like a bad model or a parser bug and is neither.

The cap was the default 4096, and this backend spends it on thinking first,
so what was left for visible content did not hold the record. The same file
already warns about this ("a 1024 cap returned nothing at all for 4 of 8
reflection prompts") -- it just had no way to raise the number from a command
line. Turning thinking off is the other lever and the worse one: it changes
what the model produces, not only how much. Measured on the architect prompt:
4096 emits ~15k output tokens across attempts in 148s and truncates about half
the time; 16000 takes 264s and parses.

`--max-tokens` and `--timeout` are declared and honoured here, which is what
this module exists for. They come in together because raising one without the
other only moves the failure: the call that now fits took 264s, and claude()'s
120s default would have timed it out three times over.

Both are opt-in, by naming a default. Nine ports already declare `--max-tokens`
themselves, each with a number it measured -- 4096, 16000, 32000 -- so
declaring it here unconditionally is an argparse conflict that takes those
entry points down at import, which is how I found them: six tests across
porous, openevolve and metasearch. A port that wants the shared flag asks for
it, the way `include_val_cap` withholds one from a port whose splits are
already frozen. The method runner asks, and keeps its measured 1024/180.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

Upstream's Architect is not a question with an answer. `agents/architect.ex` is
`use EvoGit.Agent` -- "an agent session loop template that manages a single
agent session, handling tool loops" -- with `agent_type :read_write`, which in
`agent/tools.ex` means file_read/create/write/edit, make_dir, context_read/
write/edit, run_bash, ripgrep, glob, list_dir. CONTEXT.md is written with a
tool. Nothing comes back as a structured reply and there is nothing to parse.
So is the Executor, and so is the Manager: every agent there is a session.

This port had the Executor right already (`--executor claude-code`) and the
architect as one completion per node returning {"record", "children"}. That is
the shape a reasoning model is slowest at -- one reply holding a whole record
and a whole child list, thought about once, at length -- and a truncated reply
is not a short record but an empty phase 1: one run reported "architect
designed 0 nodes, 1 replies unusable" and grew a phase 2 against no tree.

`--architect-session` designs a node the way the executor implements one: a
session in a throwaway copy, tools Read/Write/Edit/Glob/Grep, and what it did
recovered by reading the file it wrote. No Bash -- the executor gets a shell
because it must run the suite it is judged by, and an architect that can run
things is an architect that starts implementing.

One node per session, with the phase still driving the levels. Upstream's
architect spawns sub-architects and recurses inside the session; keeping the
recursion in the driver keeps phase 1's level-parallelism, keeps a session
small enough to watch, and keeps the spatial contract checkable per node.
That is a stated difference, not an oversight.

Two things fall out of the record being a file:

The design rules now live in ARCHITECT_RULES, and each path appends its own
delivery -- "reply with JSON" or "write the file". Duplicating them is how the
two paths would grow different trees.

A routing line may declare its child's size, `(12 files)`, because the
file-count threshold that refuses a child too small to be a directory has
nowhere else to live once there is no JSON. Records without counts still route.

And one defect this found, which was costing the completion path too: the
routing section was matched by an exact prefix on "## Routing Table", so
"### Routing Table", "## Routing" and "## 4. Routing Table" all opened nothing
and the node silently became a leaf. A session wrote a complete five-child
table under such a heading and the tree recorded zero children.

Measured on one fly node against a coding-plan endpoint: session 497s
uncapped, 203-242s with a 2048-token reasoning cap, same 4.6k record and the
same children. `--thinking-tokens` exposes that cap, following upstream, which
carries `reasoning_effort` per model profile beside `max_tokens` and
`concurrency` rather than fixing it. It bounds reasoning; it does not disable
it, which changes what the model produces and not only how long it takes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
`agents/manager.ex` opens with "The Manager does NOT implement features
directly" and lists five jobs: analyse, plan, delegate, **validate results**,
report completion. `agents/context_extractor.ex` is Mode A's root agent and the
reason Mode A exists -- a repository with code and no CONTEXT.md cannot be
worked on by recursive delegation, because the routing table is the map. Both
are sessions with tools upstream. Both were one completion returning JSON here.

Two of the Manager's five jobs are model calls in this port, and each now has
a session behind it.

Planning and delegation. The completion decided from the routing table and the
record alone. The routing table exists so a parent *need not* investigate the
subtree -- upstream says so -- but need not is not cannot, and the Manager is
the role upstream gives read tools to precisely so it can check.

Validating results. This is the one that was actually wrong. The completion was
handed a diff rendered and truncated at 12 000 characters; a reviewer judging a
truncated rendering of the work is not reviewing the work. `ReviewSession` runs
over the parent's state with the child's edits applied, so Read, Glob and Grep
reach the files the child really wrote.

The extractor's gap was already written down in `_extract._ask`: "Upstream it
reads what it needs with a tool; an agent here has none, so the node's own
files travel in the prompt" -- eight files of six thousand characters, and the
run before that cap invented an API surface off the file names. It reads the
code now. Its record can also carry the sections upstream lists for it and an
architect has no use for: Design Decisions, Notes for Agents, Dependencies,
Test Strategy, on upstream's own rule that a section earns its place if it
"would save an agent from re-investigating or re-discovering something".

The session machinery moves to `_session.AgentSession` and the architect sits
on it, so there is one copy of materialize / run / read-back rather than two
that drift. Tool sets follow `agent/tools.ex`: `:read` roles get Read/Glob/Grep
plus the one write a read-only agent does make (its own record), `:read_write`
adds Edit. Nobody but the executor gets Bash -- the implementer has to run the
suite it is judged by, and a design or review session that can run things is
one that starts implementing.

Deliverables are files, and lines rather than JSON: a plan of five children
whose fourth line is malformed should delegate four, not none. A verdict that
does not say ACCEPT or REJECT is not a rejection -- the child did the work, and
a reviewer that cannot speak is not evidence against it.

Still this port's standing difference, stated rather than hidden: the recursion
lives in the driver. Upstream's agents spawn subagents and recurse inside the
session; here one session does one node and the phase or the delegation policy
drives the next.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
A fly formation run reported `sessions=52 failed=43` -- 83% of implementation
episodes judged failed -- while the parent review rejected only three and the
spatial contract dropped nothing. The CLI writes a transcript per session, and
all 52 were on disk, so the cause was measurable rather than inferable:

  27 ran into the wall (>=860s of a 900s limit), 25 killed mid-tool-call
  17 had the endpoint drop the stream first (26s-830s, also mid-tool)
   8 ended on a text turn, having finished
   0 reached the 128-turn child budget -- the busiest made 97

Three defects, not a tuning problem:

* The executor was the one role built without the run's session settings. Every
  other role got `**_session_kwargs()`; this one got three arguments of its own,
  so `--timeout 600` and `--thinking-tokens 2048` reached the architect, the
  manager, the reviewer and the extractor and not the role that writes the code.
  It ran the whole domain at a 900s default nobody chose, with no reasoning cap.

* `--timeout` is the timeout on one model call (`_common` says so in its own
  help, and genesis defaults it to 120s). A session is a loop of many calls, so
  the wall is now its own flag, `--session-timeout`, printed in the run header
  beside the turn budget.

* `turns=309` was not the truth. There is no JSON after a SIGKILL, so `num_turns`
  is never read for a session that hit the wall; the transcripts hold 2607
  assistant turns. Timeouts are now counted apart from every other failure
  (`failed=43 timeout=27`), because the two have different fixes.

Work was never discarded: the port reads the worktree, not the exit status, so
a session interrupted at the wall still returns what it wrote. That was already
right and is now covered by a test.

Also, from the same transcripts: `subprocess.run(timeout=)` signals the CLI and
nothing else. A session is told to run the suite, a suite run starts servers,
and one was still listening two hours later with its working directory deleted.
Sessions now lead their own process group, killed on the way out -- after the
wall and after a clean finish alike.

Two smaller things found while reading: `--mode a --agent-sessions` raised
NameError, because the extractor read `use_sessions` and `_session_kwargs` from
above their definitions; and `ArchitectSession.strays` reported 0 unconditionally
from a helper that walked the workspace with a module its file never imported.
The session owns the workspace and deletes it, so it does the counting now.

Full suite green; six new tests, including one that asserts a session's
grandchild does not outlive the episode that started it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
Reading the executor transcripts for the wall turned up what else they carried.
A run launched from inside a Claude Code session had every episode inheriting
that session's whole situation: its MCP servers, its skill listing, its agent
types, the user's email address, and a system prompt about reviewing pull
requests and publishing artifacts. Same endpoint, same one-line prompt:

  inherited                     31,850 input tokens   6.7s   224 KB transcript
  --bare --strict-mcp-config     1,317 input tokens   2.4s    16 KB transcript

A 24x prefix on every turn of every episode, none of it about the objective and
some of it competing with it -- an executor told it is accountable for a pull
request has been given a second job. It also feeds the wall: 30k extra tokens
per turn, ~50 turns a session, 52 sessions.

Three things were shared and are now not:

* Identity. CLAUDE_CODE_SESSION_ID and its siblings are dropped, so each episode
  is its own session. Inherited, all 52 wrote transcripts named with the host's
  session id, and their TodoWrite state -- keyed by that id -- landed in the
  host's own task list, two hundred entries of "Implement src/brain package".

* State. CLAUDE_CONFIG_DIR points at one directory per run, so transcripts,
  todos and synced skills stay out of ~/.claude, which one run had left 685
  project directories in. Reported at the end of the run and deliberately not
  deleted: those transcripts are the only record of what an episode did, and
  reading 52 of them is how the wall was found.

* Context. --bare --strict-mcp-config, dropping hooks, LSP, plugin sync, commit
  attribution, auto-memory, MCP servers and CLAUDE.md auto-discovery. The
  artifact carries CONTEXT.md records the brief names, not a CLAUDE.md.

Bare mode reads credentials strictly from ANTHROPIC_API_KEY, never OAuth and
never the keychain, so it is used only when the run brought its own key: a
session billed to the local CLI's sign-in makes no API call at all with it
(measured, duration_api_ms 0), which is a worse failure than a long prompt.

The three fences are untouched. A bare session still honours
.claude/settings.local.json -- checked against the real CLI by asking one to
append to a denied path and watching it refuse, while it read and wrote the
files it was entitled to.

Full suite green (2939 tests).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
The previous commit described it as "a system prompt about reviewing pull
requests and publishing artifacts" and an executor "given a second job". That
was wrong, and the transcript's own prompt_snapshot says so: the system prompt
is 5,720 characters, and the tool schemas are 178,742.

26 of them, of which Artifact alone is 64,168, then Monitor at 14,335 and
DesignSync at 13,255. The executor is allowed six tools; Read and Bash together
are 6,887 characters of that list. Beside it ride a 13.5 KB skill listing, a
3 KB agent listing, 900 bytes of deferred tool names, and the user's email.

And one measurement that was missing, which is the whole mechanism:

  inherited                                       31,850 input tokens
  with the port's --allowedTools/--disallowedTools 31,099
  --bare --strict-mcp-config                        1,317

Permission flags do not shorten the request. They say what the session may
*call*; every other schema is sent regardless, and the port had been passing
those flags all along. Nobody instructed the episode to review a pull request --
it was handed the equipment of the session that launched it, and equipment is
priced per turn whether or not it is reachable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
…e fences are

Reading one executor's full trajectory from the isolated run turned up two
things the code was wrong about.

**--bare does not expose Write.** Neither --tools nor --allowedTools brings it
back: a bare session given `--tools Read,Write,Glob,Grep` answers "I only have a
file Read tool available". One executor spent a turn finding out --
`No such tool available: Write. Write is disabled for this session` -- and then
wrote every file through `cat > f << EOF`, while the same code without --bare
had made 1,910 successful Write calls. Edit is there and creates a file that
does not exist, which is how all seven phase-1 records in that run were written
by sessions with no shell at all. So the tool list now drops Write when bare is
on and keeps a write path that exists.

**--allowedTools is not a fence.** It is the auto-approve list; under
--permission-mode acceptEdits a session reaches for whatever built-in tool it
likes. Measured over that run's role sessions: the manager used Edit 41 times
and the reviewer 6, and Edit was in neither one's allowed tools. What does hold
is --disallowedTools -- no role session ran a shell or reached the network --
and AgentSession.run, which returns the paths the caller named and drops
everything else. The comments and the test that claimed otherwise now say this,
and the test asserts the fences that exist rather than the one that does not.

**The worktree bounds what survives, not what is seen.** Four of twelve episodes
ran `find /` and read a previous run's output from /tmp; one opened the very
_cli.py that answered the acceptance failure it had been handed to reproduce.
The brief now forbids reading outside the checkout, which is a rule and not a
wall: enforcing it takes a container, and short of that a machine should be
cleaned of earlier runs' output before starting one. Said plainly in the module
header and in the doc rather than left as an implied guarantee.

Full suite green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
The blind property was a line in a prompt. Four of twelve episodes in one run
ran `find /` and opened a previous run's output from /tmp; one of them read the
very _cli.py that answered the acceptance failure it had been handed to
reproduce. The worktree bounds what a session's work *becomes* -- an edit
outside the node is a request, a frozen write is dropped, nothing survives but
the diff -- and none of that bounds what it can read.

agentdescent already owns the fix and I had missed it. sandbox.py manages
*lifetime* and says so ("It does not manage isolation"); sandbox_container.py is
titled "A sandbox that is actually a boundary" and gives exactly what is wanted:
only the workspace visible, read-only root, no capabilities, no new privileges,
resource ceilings, network off unless the spec asks. The first look here
reported no engine and that was wrong -- dockerd was installed and simply not
running.

examples/genesis/_sandbox.py subclasses ContainerProvider and adds the mounts an
*agent* session needs that a candidate's test run does not: the claude binary's
own install (the image has no agent in it), the proxy's CA bundle (a session
that cannot verify a TLS-inspecting proxy spends its turns on certificate errors
-- a probe lost two to pip's CERTIFICATE_VERIFY_FAILED), and a per-session CLI
state directory so the transcript outlives the container. Nothing else of the
host. Asked from inside, on this machine:

  ls /home/user/agentdescent       No such file or directory
  find / -name 'algo-genesis.md'   (nothing)
  ls /tmp | wc -l                  0
  ls /work                         CONTEXT.md md.py spec src tests
  touch /etc/x                     Read-only file system
  grep CapEff /proc/self/status    CapEff: 0000000000000000

Four things this needed that the flags alone did not give:

* `network="inherit"` leaves the engine's default bridge, which is not the
  host's network. The endpoint here is reached through a proxy on the host's
  loopback, and on a bridge 127.0.0.1 is the container. The subclass adds
  --network host.
* The binary under the mount is the *resolved* path: `claude` is a symlink out
  of a node install, and exec'ing the link's own path inside is "stat: no such
  file or directory".
* Only the agent's argv[0] is rewritten. A shell probe through the same
  workspace is not the agent, and rewriting it turns `sh -c 'ls /'` into
  "Please run /login".
* The provider's own .agentdescent-mount marker is not the session's work; the
  first sandboxed diff reported it as an edit the node never made.

Also dropped MAX_THINKING_TOKENS from the inherited environment. It is not
identity but it leaks the same way: the reasoning cap the *host* session runs
under silently overrode --thinking-tokens.

Two honest limits, both stated in the doc: the network is on, so this is a
boundary against contamination and not against hostile code; and where no engine
answers the run says so once and falls back to a plain directory rather than
pretending.

Full suite green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
The default image is a name rather than a build, so a run never fails at its
first episode for want of a build context. What it does not carry is a test
runner -- and an episode is told to run the suite it is judged by, so every one
of them pays for installing it: a probe session spent two of its turns on pip,
one of them on the proxy's CERTIFICATE_VERIFY_FAILED.

--sandbox-image lets a machine that has built one point at it. Two lines and a
default that does not change.

Genesis tests green (the driver's parser is exercised by
test_the_session_wall_is_not_the_model_call_timeout); the full suite was green
on the commit before this one and this changes one flag.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
The driver hands every session role `_session_kwargs()` -- model, wall, reasoning
cap, and now the sandbox. Four of the five take `**kwargs` into AgentSession and
picked it up; ArchitectSession spells its signature out, so adding the sandbox
silently dropped it out of that contract.

The run died at phase 1 with `unexpected keyword argument 'sandbox'` *after*
printing its whole header, which is the worst shape for this failure: eighteen
lines of a working run, then a traceback, and a background task that reports
exit 0 a minute after it started.

A test now builds all five roles from one kwarg set and asserts each of them
carries it, which is the contract the driver has always assumed and never
checked.

Full suite green. The sandboxed run is past its first phase-1 node with six
containers up.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
… next start

A session container outlives its `release` when the process holding it dies. One
run was four minutes into phase 2 with eight episodes in flight when the machine
running it restarted, and afterwards all eight were still up and idling, holding
memory and a workspace mount apiece.

The engine already knows how to find them -- they carry its label and a start
time, and `ContainerProvider.reap` removes the ones past the TTL, which by
construction belong to nobody. Calling it once when the sandbox is built means
the next run cleans up after the last one, and the count goes in the run's own
report rather than being something you have to go and look for.

Housekeeping never fails a run: a reap that raises is a reap that returns zero.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
An episode is told to run the suite it is judged by. A suite run leaves a
`.pytest_cache`, and `_read_tree` skipped `.git`, `.claude` and `__pycache__`
but not that -- so the directory reached the accepted version and was
materialised into every later workspace, where the next agent reads it as if it
were the project.

Caught four hours into a live run, in the transcripts. A manager session
situated at `src/arena` opened `src/arena/.pytest_cache/`, spent seven turns
reading `.gitignore`, `CACHEDIR.TAG`, `README.md` and `v/cache/lastfailed`, and
then wrote its plan to `src/arena/.pytest_cache/.genesis/plan.md`. AgentSession
reads back the path the caller named and nothing else, so that plan was never
read: the episode cost a whole session and returned nothing. One workspace had
nested the directory twice --
`src/brain/central_complex/.pytest_cache/.pytest_cache/`.

It also shows up from outside as the thing that made the run look like it was
slowing down: new code files per check went 18, 27, 14, 10, 4 while sessions
per check went 21, 33, 33, 36, 54.

ARTIFACT_DIRS now names what a session *makes* rather than writes, and both
readers skip it -- the executor's diff and the role sessions' stray count. A
virtual environment is in the list without having been seen yet: it is the same
failure with three orders of magnitude more files.

Full suite green. The live run started before this landed and does not get it;
the next one does.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
ARTIFACT_DIRS keeps `.pytest_cache` out of the accepted version, which is where
the problem starts. It is not where it ends. In the run that turned this up the
directory got in anyway, and the managers delegated *into* it: 21 executor
episodes were situated at a cache directory, nested as deep as
`src/brain/central_complex/.pytest_cache/.pytest_cache/.pytest_cache`. Each was
an episode, a container and a session spent reading `CACHEDIR.TAG`.

So the delegation refuses it too. A node is a *source* directory; a dot
directory is never one, and neither is anything else in ARTIFACT_DIRS. Counted
as a mistaken node, which is exactly what it is -- a path read as a node when it
is not one -- beside the existing refusals for a file and for a module the path
would shadow.

Two fences on one failure because they fail differently: the first stops the
directory existing in the state at all, and the second holds even when something
else puts one there.

Full suite green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
A run's phase 1 reported `architect designed 1 nodes, deepest 0` from a record
that named six children in plain sight. The architect session was fine -- six
turns, no errors, a 3,924-character record with a full routing table. The parser
threw it away.

The brief asks for `(12 files)` and the pattern read `\d+`, so `(~32 files)` --
"about thirty-two", a perfectly reasonable thing to write -- did not match. The
count sat in an optional group that still had to match wherever it appeared, so
failing it failed the *whole line*: the child did not lose its size, it
vanished. Six lines, six vanished children, one node.

The bracket now swallows anything and the number is dug out of it afterwards, so
`(~32 files)`, `(about 32 files)`, `(32 files, maybe more)` and `(several
files)` all keep their child, and only the last loses its count.

Rare and expensive, which is the worst combination: across every run on this
machine, 7,244 routing lines carried an exact count and 49 an approximate one --
0.7%, and one of those 49 landed on a root architect and cost a whole run. A
one-node tree has no delegation, no parent review and no spatial contract in it
at all.

Full suite green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
The blind failure ended "Write your own test that reproduces this, put it
beside the code it covers, and make both pass". Every agent in a rollout is
shown that same text, and at a node it reads as an order to fix the import
here -- so eight different nodes each wrote their own _cli.py, 79 writes
between them, src/brain/olfactory/_cli.py alone 21 times. Only the one at
the repository root could ever have resolved the root import; the rest were
local imitations of a file that has to exist elsewhere, and work the parent
then had to undo.

The evidence now says it is the repository's, not necessarily yours: judge
it against what you own, write the test only if the fix belongs under your
path, and otherwise name the path in the final message and leave it alone.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
The spatial contract caps a proposal at max_files_per_diff, and the cap
rejects the diff rather than the files over the line -- so an episode one
file past it contributes nothing. At 6, counting distinct paths written
across 309 productive sessions, 11% of the episodes that did work lost all
of it; the largest was carrying 23 files. Nothing said so: the two existing
counters are about authority (wrote outside its subtree, mistook a file for
a node) and neither moves for this, so a run that lost an eighth of its work
and a run whose agents had nothing to say printed the same header.

Two fixes. The drop is counted, as discarded_diffs=N (M files) in the world
summary. And the cap stops pretending to be a trust region: the trust region
is the node's subtree, enforced edit by edit, and twelve files under one node
are not more dangerous than six. For a session executor it is now a runaway
guard at 64 -- three times the largest legitimate episode, where only a loop
that dumps a tree reaches it. 24 is where it stops binding at all. A single
completion asked for whole files keeps 6, because there six really is one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
Between "Growing the world" and the final summary the run printed nothing
at all, and for a formation run that is hours. Asking whether anything was
being accepted meant reading CLI transcripts and looking inside live
containers -- and a live worktree holds the base state PLUS whatever the
session has written since, which is not the accepted state. Two readings out
of three drawn that way were wrong.

evolve() already takes an on_round hook; the run now passes one and prints a
line per merger sweep with reward, artifact size, committed, rejected,
rollouts and elapsed. RoundInfo.reasons goes on the end because committed=0
has two causes that need opposite fixes: the gate refused the work, or the
work never reached the gate.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
A rollout is a walk of the whole Context Tree with one Claude Code session
per node, and it was a for loop. On the fly tree -- 8 nodes, deepest 2 -- a
session takes five to fifteen minutes, so a rollout took about an hour.
py-spy on a live run, 75 minutes in, showed all four workers still inside
their FIRST propose(), three levels deep in nested _episode frames. The
accepted state held three files.

That read as an acceptance problem and was not one: the parent gate accepts
a tie, four of the five reviews that had run said ACCEPT, and the staleness
policy keeps a card whose reward is merely unchanged. Nothing was being
rejected. Nothing had finished. And --workers counts whole tree walks, so
four workers were four independent descents of the same eight nodes -- two
of them on src/brain/navigation at the same moment, doing the same node's
work, of which one result could survive.

Siblings now run concurrently and are folded in delegations order, so the
proposal a rollout returns is unchanged. The spatial contract is what makes
that safe: a child writes only under its own subtree, so two siblings cannot
touch the same path. --node-workers says how wide one level may get,
--max-sessions bounds what the endpoint actually sees, and the bound lives
in run_cli rather than in the pools -- sessions are the scarce thing, and
they never nest, so one semaphore there covers the whole tree with no way
for a parent to deadlock on its own children.

Every counter the port reports is now incremented under a lock. x += 1 is a
read and a write with a bytecode boundary between them, and these numbers
are the port's evidence.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
@Birfy Birfy changed the title Genesis: blind mode, a product domain, and fourteen defects a real run found Genesis: blind mode, a product domain, and eighteen defects a real run found Sep 16, 2026
… of them

--bare removes the inherited tool schemas and a signed-in run cannot use it:
it sets CLAUDE_CODE_SIMPLE=1, and that same switch makes the CLI refuse to
read OAuth (measured: duration_api_ms 0 and an authentication error). The two
cannot be separated, so such a run got the whole prompt -- and the whole
prompt is mostly schemas. Asked to list its tools, one of these sessions names
forty-two of them: Artifact, CronCreate, DesignSync, PushNotification,
Workflow, ShowOnboardingRolePicker and the rest, none reachable from a
throwaway worktree, all priced every turn.

--allowedTools does not help; it auto-approves, and every other schema is sent
regardless. What decides which schemas exist is which tools are defined, and
--agents defines its own. Measured on one endpoint with the same one-line
prompt: 24 550 tokens of context per turn with --allowedTools, 6 522 with an
--agents entry declaring five tools, 8 042 for the port's own executor command
(three more tools, Bash among them), against 4 868 under --bare. Two thirds of
the tax, gone, with the sign-in intact.

lean_agent_flags returns nothing when --bare is already in play, so a keyed run
still takes the better path and the declaration is only what a sign-in gets
instead. The remaining gap to --bare is the system prompt, which --agents
cannot touch.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
…andbox

A key is a string in the environment and crosses into a container with it,
which is why a keyed run is isolated and authenticated at once. A sign-in is
not a string: the CLI reaches the endpoint through the host's session ingress,
which the container does not have and which this module does not put there.

So every containerised session answered "Not logged in · Please run /login",
and the run found out one episode at a time -- sessions=4 failed=4 edits=0,
architect designed 0 nodes, a whole run spent on a condition that was knowable
before the first episode started.

SessionSandbox now reports it beside the two reasons it already gave, a missing
container engine and an unmountable toolchain, and degrades to a plain
directory with the reason printed. The remedy it names is an API key, which is
also the arrangement with the smallest prompt, since a key is what --bare
needs. Running signed-in without the sandbox is the other option and is worth
reading carefully: the session is then on the host as the host's user and
reaches everything the sign-in reaches anyway, so what that trade gives up is
the isolation, not the exposure.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
A rollout descends the entire Context Tree -- the root delegates, every leaf
writes, the parents fold, and one proposal carries all of it. Two bounds stood
between that and the accepted version. SpatialContract.max_files_per_diff is
the one raised to 64 for a session executor; RecursiveDelegation.max_edits sits
upstream of it in _bound and was 4, a trust region sized for a single
completion proposing a file or two. Only the tighter one ever applied, so
raising the other changed nothing.

Measured on an 8-node fly tree: the sessions wrote 153 implementation files,
every sweep committed exactly 4, and after three rollouts the accepted state
held nine records and six Python files. src/training wrote 26 of them,
src/brain/circuits 24, src/arena 22, src/frontend 21 -- and the version grew by
four a round.

Both now read one named value. What _bound does within the bound is unchanged
and was not a defect: a node-creating record is trimmed last, because nothing
else in the accepted version says the node exists and the source file it would
be dropped for is re-proposable next round; work next; routine upkeep first. A
test pins that order, and another pins the two caps to one name so they cannot
drift apart again.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
@Birfy Birfy self-assigned this Sep 16, 2026
A proposal crosses five bounds between the session that writes a file and the
state that keeps it. Two of them were this port's and were raised together. Two
are the engine's -- `trust_region_ops` and `trust_region_chars` -- and nothing
here had ever passed an `agg_config`, so the tightest cap on the path was a
default of six ops, sized for a rule table and applied to a proposal carrying an
eight-node repository.

Measured: one rollout, thirteen sessions, 680 turns, forty minutes, merged
`committed=0 rejected=1 [oversized=1]` with `discarded_diffs=0` and
`truncated_edits=0` -- the port's own bounds passed it through cleanly and the
engine rejected it behind them. Nine CONTEXT.md records landed and not one line
of implementation. The largest single file any session wrote was 12,813 chars,
so only the op count was ever binding.

`engine_bounds(strategy)` derives the engine's two from the strategy that
carries the cap, so there is one gate and it is the one that counts what it
drops. `batch_trigger` and `max_wait_rounds` are carried across because
`evolve()` builds `AggregatorConfig(batch_trigger=2, max_wait_rounds=1)` when
passed nothing, not the dataclass defaults, and a config passed in replaces that
object whole -- a run widening its trust region would otherwise have doubled its
batch trigger in silence.

Three tests, each verified to fail against the old behaviour: the caps all read
one number, a diff at the full cap clears the engine's real trust-region
predicate, and the batching numbers are read back out of a live `evolve()`
rather than repeated as literals.

`agentdescent/` is unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016CjFJhQ2Wxq9Vr3Fu38KrW
@Birfy
Birfy marked this pull request as draft October 10, 2026 13:09

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants