v4.12.0 campaign: sessions, resources, shell execution and tool usability - #635
Merged
Merged
Conversation
Follow up 3f92cc83 for #632: control synthetic evidence and action leases from fixture creation, including scope overrides and guardian refresh. Exercise the actual focus cases with 300ms evidence-processing delays under coverage, while retaining real acquisition timeouts, freshness boundary tests, and interruption/no-replay semantics.
Calmingstorm
marked this pull request as ready for review
September 30, 2026 19:32
This was referenced Oct 1, 2026
Closed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Approved v4.12.0 campaign
Implements all fifteen items in the Aaron-approved work order on
campaign/v4.12.0, based on latest mastera93348f0. Existing configs and persisted records remain loadable. No merge, deployment, restart, release dispatch, live-install testing, issue closure or attribution trailers.LimitNOFILE=65536retained and parsed-unit regression added; actual process open-file utilization health component warns strictly above 70%, rendered by Health page.run_command,run_command_multi, local process starts, command validation checks and skillrun_on_host; shared runners/internal helpers remain/bin/shwith unchanged output. Configurable `autoRead-only automation inventory for B3
Inspected existing persisted schedules and ten registered skill templates without executing their commands or invoking the skills. Credential values are excluded.
server:/opt/eve-intel-toy/.venv/bin/eve-stream-report --database /var/lib/eve-intel/eve-stream.sqlite3 --hours 24 --discord-reportcurseforge_manage,dynamic_odyssey_releasecyberpower_upspwrstatpaths with2>&1and `amp_managess/sed/tail/find/sort/headpipelinesminecraft_controlmc_modpack_updaterun_on_hostliteralsNo dash-specific constructs or Bash-only syntax were found in these persisted templates. Dynamically supplied arguments are not an exhaustive runtime inventory. Upgrade/rollback guidance is in
docs/command-shell-upgrade.mdand CHANGELOG:tools.command_shell: shis the immediate rollback.B3 lifecycle findings and qualification boundary
The sh/bash matrix exposed a rapid-exit settlement-channel race. The worker now retains an already-verified-empty ownership channel until clean settlement is consumed and acknowledged; the acknowledgement is not cleanup evidence. Normal foreground results return at leader exit plus output EOF, without waiting for or terminating surviving descendants. Delayed/wrong ACK and disconnect tests preserve fail-closed restart vetoes and ownership rules.
Governor coverage evaluates command strings only. It never executes a destructive or governor-blocked fixture. Unsupported/over-budget active brace expansion fails closed instead of dropping dangerous alternatives.
Aaron's scope correction is implemented: code-built
read_file,apply_patch, script wrappers, HTTP probes, non-command validation probes and every default internal caller use/bin/shregardless of configuration, without annotations. Actual kernel executable/argv tests underautoandbashverify byte-identical internal output, and independently traced process execution confirms opted-in Bash versus internal POSIX dispatch.The requested comparative live soak remains a separate, post-branch-deployment operator qualification. It was not run during this development task. No operational commands were shadow-run.
Verification
Final head:
57c0426a5a391b822619b1174ed9d55b0437177d. All applicable hosted checks green; deploy skipped.npm run check, WebUI build and Pages build: passed. Combined regeneratedui/distcommitted.Earlier gates exposed a stale worker ACK fixture hang, internal-output annotation corruption, stale log/subprocess mock contracts, generated-reference drift, prompt-size regression and raw transport recovery classification regression. These were corrected with real behavior regressions, not rerun unchanged into green. Aaron's subsequent scope correction was implemented before final verification. Three coverage misses were covered with additional tests; the baseline was not edited.
Warnings remain visible, including aiohttp AppKey and audioop deprecation warnings, Vue fixture lifecycle warnings and the existing large Vite chunk warning. Twenty-seven skipped tests remain environment/optional-fixture dependent. Production freshness/cleanup assertions were not weakened. The branch qualification soak remains explicitly unperformed.
Round 1 archived verification evidence
The following evidence was moved out of the published developer docs in round 2. The 1ms startup cadence and the system-prompt shell reminder described here were subsequently removed in round 2. Its measurements are historical, not qualification of the new head. The standalone compatibility and latency scripts were removed; collected regression tests remain.
PR 635 round 1: compatibility evidence
All probes ran in development trees, using temporary files and privately owned
sleep/Python children. No live-install tests, services, deployments or agents.
Baseline: v4.11.0 master
a93348f004.Foreground lifetime and output
tests/test_pr635_round1.pycoverssh,auto, andbashthroughrun_command,run_command_multi, skillrun_on_host, and streaming:nohup, setsid, and fork/exit-parent children survive normal foreground return.
Separate tests verify timeout, cancellation and shutdown reap those same fixture
types, using exact supervisor ownership. No numeric-PID cleanup is used.
tests/pr635_compat_probe.pywas run against master and this branch in eachmode. Every recorded
(code, text, recovery category)matched master exactly:successful command/multi/skill output; missing and unreadable files; apply_patch
host failure; HTTP transport timeout; script timeout and nonzero exit. Every
normal-completion background child was still alive four seconds after return,
and was subsequently reaped by the private supervisor shutdown.
Settlement and latency
The existing ACK protects against the observed rapid-exit socket-close/write
race, but is no longer on the foreground return path. The monitor's ACK drain
and the empty worker's ACK wait are each bounded at two seconds. Delayed and
missing ACK tests verify respectively that foreground success does not wait and
that an already-empty worker cannot hang indefinitely.
An initial comparison exposed a 20ms startup polling quantization with bash.
The worker now uses a 1ms control-poll interval only during its first 100ms;
long-lived jobs keep the existing 20ms interval. Ownership and teardown are
unchanged.
Final interleaved samples (
tests/pr635_latency_probe.py, 60 measurements eachafter five warm-ups), milliseconds:
No measured regression; these are workstation samples, not a universal latency
guarantee. In particular, ACK completion is asynchronous, not charged to the
foreground call.
Remaining review items
Effective-shell text was removed from command results and process start/poll/list
presentation. Dynamic contracts, durable process records and process API fields
remain. Internal transports and recovery prefix gates retain the legacy contract.
Every pre-existing system-prompt line is identical to master, with only the shell
reminder added and the two size tests raised to 5400 characters.
Brace-range tests are classification-only; no dangerous classifier input is
executed. Session rollback tests restore one valid and one future-dated record,
discard only the future record, then persist and restore a new valid session.
Round 2 changes and safety boundary
All work done directly, without agents. No merge, deploy, restart, live-install test, pipeline dispatch, issue closure or attribution trailers. Sealed cross-distro development qualification completed below; Claude's independent acceptance rerun remains as specified in R2-0.
Round 2 local verification
Head
e7686ad697b95f7d00e7bba80bcdca0bedcd82c5.npm run check, WebUI build and generated tool/API reference parity passed.R2-0 sealed cross-distro qualification
Built isolated disposable Python 3.12 images for Alpine 3.24.2, Ubuntu 24.04.5 LTS and Debian 12. Runtime containers had no network, all capabilities dropped, no-new-privileges, read-only root and source mounts, private tmpfs fixture directories, no Docker socket or desktop/runtime-config mounts. Executed only the supplied harmless seven-patch repro. Dependency installation happened during image construction, not during the sealed checks.
Concurrent stress: 240 iterations × 7 calls × 3 distros = 5,040 calls, all six valid patches succeeded, every seventh mismatch refused exactly as expected, zero ownership-lost results and zero settlement warnings. Additional alternating-order cost qualification added 2,520 branch calls and 2,520 master calls, again zero ownership failures/warnings. No R2-1 classifier string entered these probes.
Two opposite-order samples of 420 calls per tree/distro, whole-harness elapsed seconds (master
a93348f0, branche7686ad6):Observed cost is essentially parity: Alpine and Debian slightly lower; Ubuntu +0.526ms/call (+0.68%) in these finite samples. Do not claim a strict cost win on Ubuntu or treat noise as proof. The requested zero-failure stress threshold is met; Claude retains independent review of the cost acceptance. All containers auto-removed and the temporary master worktree removed after measurement. Logs retained under
/tmp/odin-pr635-r2-distro-{stress,cost}.logfor review, not published developer docs.Extended Ubuntu check, opposite-order 840-call samples: master 64.298640s / 64.347703s, branch 64.557576s / 64.667603s. This confirms a small measurable +0.345ms/call (+0.45%) whole-harness cost, not a strict master-or-better result. All additional 1,680 branch and 1,680 master calls again had zero ownership errors or warnings. Total sealed branch calls: 9,240, including the 5,040 concurrent stress calls. Therefore the zero-failure requirement passes, but the literal Ubuntu cost criterion remains unmet in these samples and is explicitly reported rather than rounded into a pass.
Round 3: R3-1 event-driven exit detection
Head
aa1f592315cfb55578c3ad212f8933e14d1f497f. The worker registers a dedicated leader pidfd and each verified descendant's existing pidfd in its selector. Exit readiness wakes the existing ownership loop immediately. Readiness consumes only the watch, not the ownership descriptor, avoiding level-triggered exit spinning; reporting/reaping unregisters and closes descriptors. Disconnected control channels still wait on process readiness. Idle fallback remains 20ms, with no startup fast-poll window. Discovery, vanished-process tolerance and reaping decisions are unchanged.Regression coverage includes leader/descendant registration, mixed control/exit events, one-shot watch retirement, reap after consumed readiness, registration-failure cleanup, and real harmless child exit wakeups with both connected and disconnected control sockets. Existing foreground-background lifetime, ACK, cancellation and cleanup tests pass.
Sequential latency
The supplied
latency_ab.pywas run unchanged first, then repeated with the same measurement body plus result assertions and an explicit private supervisor shutdown barrier. The original script cancels outstanding settlement monitors on event-loop shutdown, producing teardown warnings on master and branch; the explicit barrier removes those harness warnings without changing measured command timings.100 sequential calls per sample after five warmups, opposite orders, master
a93348f0:Observed auto p50 is within 0.8ms of master and 0.6ms of sh, meeting the 2ms local acceptance bound. This finite desktop measurement is not a universal latency guarantee. Evidence:
/tmp/odin-pr635-r3-latency{,-clean}.log.Sealed cross-distro stress
Alpine 3.24.2, Ubuntu 24.04.5 and Debian 12: 5,040 seven-patch calls plus 5,040 sequential opted-in auto
run_command truecalls per distro, 30,240 total calls. Every valid patch succeeded, each intentional context mismatch refused, every short command succeeded; zero ownership-lost results and zero settlement warnings in all six runs. Each harness explicitly settled its own supervisors before exiting. Runtime containers had no network, no capabilities, no-new-privileges, read-only root/source mounts and private fixture tmpfs; no live config, Docker socket or desktop mounts. All auto-removed. Evidence:/tmp/odin-pr635-r3-stress.logand harmless harnesses/tmp/odin-pr635-r3-{stress.sh,patch-repro.py,command-stress.py}.Gates
Final-head hosted gates are all green, deploy skipped: https://github.com/Calmingstorm/Odin/actions/runs/36811534291. Plain and coverage suites each 23,534 passed, 27 skipped; 94.6% reported coverage, zero ratchet findings, unchanged baseline. Local final full coverage run: 23,535 passed, 26 skipped, 94.6%, zero findings. Targeted regression tests, lint/type ratchets and config classification passed; browser guard (six real-browser tests),
npm run check, WebUI build and Pages build passed. Pages initially lacked its declared local dependency;npm --prefix site ciinstalled the lockfile dependencies and the build passed. Existing deprecation, aiohttp/AppKey, async teardown/unraisable, Vue lifecycle and large-chunk warnings remain visible; site dependency installation also reports five audit advisories (one low, three moderate, one high), not changed in this scoped round.The first full local/hosted suites found two identity-reuse fixture failures: the old test bypassed
Worker.__init__and had no selector. The test-only follow-up initializes the real worker/control socket and registers its exact owned pidfd, asserts watch retirement and unchanged start-ID safety, and closes all fixture resources. The 147-test shell matrix/round-2/pidfd rerun passed. No production exception was hidden and no cleanup assertion was weakened.No agents, merge, deployment, restart, release dispatch, live-install testing, issue closure or attribution trailers.