Skip to content

v4.12.0 campaign: sessions, resources, shell execution and tool usability - #635

Merged
Calmingstorm merged 32 commits into
masterfrom
campaign/v4.12.0
Oct 1, 2026
Merged

Calmingstorm merged 32 commits into
masterfrom
campaign/v4.12.0

Conversation

@Calmingstorm

@Calmingstorm Calmingstorm commented Sep 30, 2026 •

Copy link
Copy Markdown
Owner

Approved v4.12.0 campaign

Implements all fifteen items in the Aaron-approved work order on campaign/v4.12.0, based on latest master a93348f0. Existing configs and persisted records remain loadable. No merge, deployment, restart, release dispatch, live-install testing, issue closure or attribution trailers.

Item Implementation and regression evidence
A1 / #633 Authenticated request and WebSocket expiry returns to sign-in once, stops polling/socket reconnects, preserves 403 and newer sessions; isolated real Chromium/server-session invalidation tests.
A2 Explicit persistent-login opt-in; atomic 0600 hashed-ID store and private HMAC key, current issuing-credential/policy revalidation, wall-clock idle/30-day expiry, coalesced activity, durable logout and corrupt-store failure fencing; real HTTP and restored WebSocket tests.
A3 Skill tests run under signed-in task-local identity, tier, host/default-host and credential/tool grants; real skill-manager/executor parity tests.
A4 Backward bounded-block tail reads off the event loop, last 50 complete records, partial-record cursor and rotation handling; large-file bounded-read/memory and rotation tests.
B1 / #634 Finished and restored outputs retain paths rather than descriptors; bounded reads open/close transiently; actual descriptor, capture, cursor replay and expiry tests.
B2 Existing packaged LimitNOFILE=65536 retained and parsed-unit regression added; actual process open-file utilization health component warns strictly above 70%, rendered by Health page.
B3 Explicit raw-command opt-in only: run_command, run_command_multi, local process starts, command validation checks and skill run_on_host; shared runners/internal helpers remain /bin/sh with unchanged output. Configurable `auto
C1 / #630, #631 Skill states and agent model/effort provenance listed; complete executable skill module contract/example and parity pins.
C2 Search timeouts include duration; empty exception strings use class names.
C3 Per-call browser selector wait defaults to 10 seconds for omitted/zero/blank, capped by classified config leaf at maximum 60 seconds.
C4 Noon, following-day midnight, AM/PM dayparts compose with day words/zones; conflicting clocks rejected; exact-instant and legacy grammar regressions.
C5 Nullable reasoning tokens across both provider response paths, incomplete agent and direct/tool-loop chat records, transactional usage migration, API/UI unknown-versus-zero display; actual wire bytes unchanged.
D1 Host-access removal audits actor, target and locked prior committed entry with signed-chain tests; 404 writes nothing.
D2 Same iteration prefix is added once; different iteration remains distinct.
D3 / #632 Fixture-owned native evidence time, with delayed setup and exact freshness-boundary/no-replay tests; production 250 ms freshness bound unchanged.

Read-only automation inventory for B3

Inspected existing persisted schedules and ten registered skill templates without executing their commands or invoking the skills. Credential values are excluded.

Persisted automation Command/template and scope Shell finding
Anniversary and birthday schedules Two reminders, no command No shell execution.
Daily New Eden Killstream schedule Remote server: /opt/eve-intel-toy/.venv/bin/eve-stream-report --database /var/lib/eve-intel/eve-stream.sqlite3 --hours 24 --discord-report Simple command; remote foreground remains unchanged.
Persisted workflow steps None present None to migrate.
curseforge_manage, dynamic_odyssey_release Localhost sudo helper calls with shell-quoted assembled arguments No dash-specific syntax identified.
cyberpower_ups Local/configurable pwrstat paths with 2>&1 and `
amp_manage Local Python heredoc/socket probe and SSH tunnel; remote sudo AMP and ss/sed/tail/find/sort/head pipelines POSIX templates; remote unchanged.
minecraft_control Local/configurable curl with quoted arguments No dash-specific syntax identified.
mc_modpack_update Remote/configurable CLI with quoted arguments Remote unchanged.
Cloudflare, photo editing, Linode, webcam skills No run_on_host literals No affected shell templates.

No dash-specific constructs or Bash-only syntax were found in these persisted templates. Dynamically supplied arguments are not an exhaustive runtime inventory. Upgrade/rollback guidance is in docs/command-shell-upgrade.md and CHANGELOG: tools.command_shell: sh is the immediate rollback.

B3 lifecycle findings and qualification boundary

The sh/bash matrix exposed a rapid-exit settlement-channel race. The worker now retains an already-verified-empty ownership channel until clean settlement is consumed and acknowledged; the acknowledgement is not cleanup evidence. Normal foreground results return at leader exit plus output EOF, without waiting for or terminating surviving descendants. Delayed/wrong ACK and disconnect tests preserve fail-closed restart vetoes and ownership rules.

Governor coverage evaluates command strings only. It never executes a destructive or governor-blocked fixture. Unsupported/over-budget active brace expansion fails closed instead of dropping dangerous alternatives.

Aaron's scope correction is implemented: code-built read_file, apply_patch, script wrappers, HTTP probes, non-command validation probes and every default internal caller use /bin/sh regardless of configuration, without annotations. Actual kernel executable/argv tests under auto and bash verify byte-identical internal output, and independently traced process execution confirms opted-in Bash versus internal POSIX dispatch.

The requested comparative live soak remains a separate, post-branch-deployment operator qualification. It was not run during this development task. No operational commands were shadow-run.

Verification

Final head: 57c0426a5a391b822619b1174ed9d55b0437177d. All applicable hosted checks green; deploy skipped.

  • Full plain suite: 23,348 passed, 27 skipped, zero failures.
  • Full instrumented suite: 23,348 passed, 27 skipped, zero failures; 94.6% reported coverage and 0 coverage-gate findings.
  • Lint and type ratchets: 0 new findings; configuration classification: 312 leaves, 0 findings.
  • Browser network guard, npm run check, WebUI build and Pages build: passed. Combined regenerated ui/dist committed.
  • Final hosted full-suite evidence: Tests run 36784510435.

Earlier gates exposed a stale worker ACK fixture hang, internal-output annotation corruption, stale log/subprocess mock contracts, generated-reference drift, prompt-size regression and raw transport recovery classification regression. These were corrected with real behavior regressions, not rerun unchanged into green. Aaron's subsequent scope correction was implemented before final verification. Three coverage misses were covered with additional tests; the baseline was not edited.

Warnings remain visible, including aiohttp AppKey and audioop deprecation warnings, Vue fixture lifecycle warnings and the existing large Vite chunk warning. Twenty-seven skipped tests remain environment/optional-fixture dependent. Production freshness/cleanup assertions were not weakened. The branch qualification soak remains explicitly unperformed.

Round 1 archived verification evidence

The following evidence was moved out of the published developer docs in round 2. The 1ms startup cadence and the system-prompt shell reminder described here were subsequently removed in round 2. Its measurements are historical, not qualification of the new head. The standalone compatibility and latency scripts were removed; collected regression tests remain.

PR 635 round 1: compatibility evidence

All probes ran in development trees, using temporary files and privately owned
sleep/Python children. No live-install tests, services, deployments or agents.
Baseline: v4.11.0 master a93348f004.

Foreground lifetime and output

tests/test_pr635_round1.py covers sh, auto, and bash through
run_command, run_command_multi, skill run_on_host, and streaming:
nohup, setsid, and fork/exit-parent children survive normal foreground return.
Separate tests verify timeout, cancellation and shutdown reap those same fixture
types, using exact supervisor ownership. No numeric-PID cleanup is used.

tests/pr635_compat_probe.py was run against master and this branch in each
mode. Every recorded (code, text, recovery category) matched master exactly:
successful command/multi/skill output; missing and unreadable files; apply_patch
host failure; HTTP transport timeout; script timeout and nonzero exit. Every
normal-completion background child was still alive four seconds after return,
and was subsequently reaped by the private supervisor shutdown.

Settlement and latency

The existing ACK protects against the observed rapid-exit socket-close/write
race, but is no longer on the foreground return path. The monitor's ACK drain
and the empty worker's ACK wait are each bounded at two seconds. Delayed and
missing ACK tests verify respectively that foreground success does not wait and
that an already-empty worker cannot hang indefinitely.

An initial comparison exposed a 20ms startup polling quantization with bash.
The worker now uses a 1ms control-poll interval only during its first 100ms;
long-lived jobs keep the existing 20ms interval. Ownership and teardown are
unchanged.

Final interleaved samples (tests/pr635_latency_probe.py, 60 measurements each
after five warm-ups), milliseconds:

Tree/mode Median Mean Standard deviation
master before sh 64.01 69.83 13.06
branch sh 61.35 63.09 11.75
master before auto 66.48 72.78 13.01
branch auto 62.82 63.89 4.04
master before bash 64.95 71.85 15.77
branch bash 62.87 64.21 9.39
master after bash 69.94 75.62 15.06

No measured regression; these are workstation samples, not a universal latency
guarantee. In particular, ACK completion is asynchronous, not charged to the
foreground call.

Remaining review items

Effective-shell text was removed from command results and process start/poll/list
presentation. Dynamic contracts, durable process records and process API fields
remain. Internal transports and recovery prefix gates retain the legacy contract.
Every pre-existing system-prompt line is identical to master, with only the shell
reminder added and the two size tests raised to 5400 characters.

Brace-range tests are classification-only; no dangerous classifier input is
executed. Session rollback tests restore one valid and one future-dated record,
discard only the future record, then persist and restore a new valid session.

Round 2 changes and safety boundary

  • R2-0: descendant discovery treats ENOENT/ESRCH and proven vanished processes as an incomplete scan, without vetoing restart or reporting ownership loss; the next iteration rescans. Real failures retain fail-closed error/termination. Removed the 1ms startup cadence. Short internal patch sequence and errno/race fixtures added.
  • R2-1: pure-Python ANSI-C decoding contributes classification alongside raw text; remote downloader process substitution feeding bash/sh/source/dot receives piped-script risk. Every dangerous input remains exclusively in classification-only tests. No dangerous input was sent to an execution backend, even a mocked one.
  • R2-2: dynamic contracts replace earlier shell decoration; repeated application and served chat/agent/autonomous-loop catalogs are tested.
  • R2-3: explicit missing bash gives a plain setting-specific command-not-executed refusal through all four raw-command routes, before any spawn.
  • R2-4: CHANGELOG and upgrade guide name echo/brace/ANSI-C/glob-locale/builtin-status/error-wording differences and sh rollback.
  • R2-5: truthful command-only validation timeout verdict; non-command legacy verdicts retained.
  • R2-6: shell selection stays in process records/API, not poll retention JSON; only failed/killed jobs add informative termination/cleanup fields.
  • R2-7: historical verification evidence moved here; the published evidence doc and two uncollected scripts removed.
  • Aaron's added decision: system_prompt.py is byte-identical to a93348f (SHA-256 69f65608ece50596ace38862bf5c93377ea24f92424a91bb2935050936c26194), both size pins restored to <5000, reminder assertions removed, fixed-input prompt equality and effective-shell dynamic contracts tested.

All work done directly, without agents. No merge, deploy, restart, live-install test, pipeline dispatch, issue closure or attribution trailers. Sealed cross-distro development qualification completed below; Claude's independent acceptance rerun remains as specified in R2-0.

Round 2 local verification

Head e7686ad697b95f7d00e7bba80bcdca0bedcd82c5.

  • All applicable final-head hosted checks green; deploy skipped. Tests run: https://github.com/Calmingstorm/Odin/actions/runs/36804949825. Both hosted full suites: 23,520 passed, 27 skipped; coverage 94.6%, zero findings. Pages/WebUI builds also green.
  • Final production-code plain suite: 23,520 passed, 26 skipped. The subsequent test-only commit adds one malformed-discovery regression, which passed plain and instrumented targeted runs.
  • Final full instrumented suite at the final head: 23,521 passed, 26 skipped; 94.6% reported coverage; zero coverage ratchet findings. Baseline unchanged.
  • Lint/type ratchets and config classification: zero new findings; 312 config leaves classified. npm run check, WebUI build and generated tool/API reference parity passed.
  • Harmless external patch repro: 5,040 calls on the final production code across local sh/auto/bash with zero ownership-lost results and zero settlement warnings. Seven-patch collected fixture also passes. Additional sealed cross-distro qualification is below.
  • Interleaved local cost comparison, 420 calls each: master-before 70.29s, branch 64.48s, master-after 54.49s under concurrent verification load. Branch lies within the master samples; this is not sufficient to claim a universal cost improvement. Claude's sealed cross-distro cost/stress acceptance remains outstanding.
  • The first paired full runs exposed an old missing-bash wording assertion, corrected to the new explicit refusal. A later instrumented run exposed three uncovered malformed-discovery error lines; the final test-only commit covers them. No unchanged retry was used to manufacture green.
  • Warnings remain visible: aiohttp AppKey warnings, audioop deprecation, asynchronous teardown/unraisable warnings, Vue fixture lifecycle warnings and the large Vite chunk warning. Task-created fixture directories and detached baseline worktree were removed.

R2-0 sealed cross-distro qualification

Built isolated disposable Python 3.12 images for Alpine 3.24.2, Ubuntu 24.04.5 LTS and Debian 12. Runtime containers had no network, all capabilities dropped, no-new-privileges, read-only root and source mounts, private tmpfs fixture directories, no Docker socket or desktop/runtime-config mounts. Executed only the supplied harmless seven-patch repro. Dependency installation happened during image construction, not during the sealed checks.

Concurrent stress: 240 iterations × 7 calls × 3 distros = 5,040 calls, all six valid patches succeeded, every seventh mismatch refused exactly as expected, zero ownership-lost results and zero settlement warnings. Additional alternating-order cost qualification added 2,520 branch calls and 2,520 master calls, again zero ownership failures/warnings. No R2-1 classifier string entered these probes.

Two opposite-order samples of 420 calls per tree/distro, whole-harness elapsed seconds (master a93348f0, branch e7686ad6):

Distro Master samples Branch samples Mean milliseconds/call master / branch
Alpine 3.24.2 48.110925, 48.006407 48.026026, 48.060206 114.425 / 114.388
Ubuntu 24.04.5 32.406386, 32.680109 32.412302, 33.116025 77.484 / 78.010
Debian 12 42.884287, 42.806526 42.787252, 42.837005 102.013 / 101.934

Observed cost is essentially parity: Alpine and Debian slightly lower; Ubuntu +0.526ms/call (+0.68%) in these finite samples. Do not claim a strict cost win on Ubuntu or treat noise as proof. The requested zero-failure stress threshold is met; Claude retains independent review of the cost acceptance. All containers auto-removed and the temporary master worktree removed after measurement. Logs retained under /tmp/odin-pr635-r2-distro-{stress,cost}.log for review, not published developer docs.

Extended Ubuntu check, opposite-order 840-call samples: master 64.298640s / 64.347703s, branch 64.557576s / 64.667603s. This confirms a small measurable +0.345ms/call (+0.45%) whole-harness cost, not a strict master-or-better result. All additional 1,680 branch and 1,680 master calls again had zero ownership errors or warnings. Total sealed branch calls: 9,240, including the 5,040 concurrent stress calls. Therefore the zero-failure requirement passes, but the literal Ubuntu cost criterion remains unmet in these samples and is explicitly reported rather than rounded into a pass.

Round 3: R3-1 event-driven exit detection

Head aa1f592315cfb55578c3ad212f8933e14d1f497f. The worker registers a dedicated leader pidfd and each verified descendant's existing pidfd in its selector. Exit readiness wakes the existing ownership loop immediately. Readiness consumes only the watch, not the ownership descriptor, avoiding level-triggered exit spinning; reporting/reaping unregisters and closes descriptors. Disconnected control channels still wait on process readiness. Idle fallback remains 20ms, with no startup fast-poll window. Discovery, vanished-process tolerance and reaping decisions are unchanged.

Regression coverage includes leader/descendant registration, mixed control/exit events, one-shot watch retirement, reap after consumed readiness, registration-failure cleanup, and real harmless child exit wakeups with both connected and disconnected control sockets. Existing foreground-background lifetime, ACK, cancellation and cleanup tests pass.

Sequential latency

The supplied latency_ab.py was run unchanged first, then repeated with the same measurement body plus result assertions and an explicit private supervisor shutdown barrier. The original script cancels outstanding settlement monitors on event-loop shutdown, producing teardown warnings on master and branch; the explicit barrier removes those harness warnings without changing measured command timings.

100 sequential calls per sample after five warmups, opposite orders, master a93348f0:

Mode p50 sample 1 / sample 2 Mean sample 1 / sample 2
Master 36.9 / 37.1ms 40.6 / 42.4ms
Branch sh 37.1 / 37.3ms 37.1 / 37.3ms
Branch auto 37.5 / 37.7ms 37.5 / 37.8ms

Observed auto p50 is within 0.8ms of master and 0.6ms of sh, meeting the 2ms local acceptance bound. This finite desktop measurement is not a universal latency guarantee. Evidence: /tmp/odin-pr635-r3-latency{,-clean}.log.

Sealed cross-distro stress

Alpine 3.24.2, Ubuntu 24.04.5 and Debian 12: 5,040 seven-patch calls plus 5,040 sequential opted-in auto run_command true calls per distro, 30,240 total calls. Every valid patch succeeded, each intentional context mismatch refused, every short command succeeded; zero ownership-lost results and zero settlement warnings in all six runs. Each harness explicitly settled its own supervisors before exiting. Runtime containers had no network, no capabilities, no-new-privileges, read-only root/source mounts and private fixture tmpfs; no live config, Docker socket or desktop mounts. All auto-removed. Evidence: /tmp/odin-pr635-r3-stress.log and harmless harnesses /tmp/odin-pr635-r3-{stress.sh,patch-repro.py,command-stress.py}.

Gates

Final-head hosted gates are all green, deploy skipped: https://github.com/Calmingstorm/Odin/actions/runs/36811534291. Plain and coverage suites each 23,534 passed, 27 skipped; 94.6% reported coverage, zero ratchet findings, unchanged baseline. Local final full coverage run: 23,535 passed, 26 skipped, 94.6%, zero findings. Targeted regression tests, lint/type ratchets and config classification passed; browser guard (six real-browser tests), npm run check, WebUI build and Pages build passed. Pages initially lacked its declared local dependency; npm --prefix site ci installed the lockfile dependencies and the build passed. Existing deprecation, aiohttp/AppKey, async teardown/unraisable, Vue lifecycle and large-chunk warnings remain visible; site dependency installation also reports five audit advisories (one low, three moderate, one high), not changed in this scoped round.

The first full local/hosted suites found two identity-reuse fixture failures: the old test bypassed Worker.__init__ and had no selector. The test-only follow-up initializes the real worker/control socket and registers its exact owned pidfd, asserts watch retirement and unchanged start-ID safety, and closes all fixture resources. The 147-test shell matrix/round-2/pidfd rerun passed. No production exception was hidden and no cleanup assertion was weakened.

No agents, merge, deployment, restart, release dispatch, live-install testing, issue closure or attribution trailers.

@Calmingstorm
Calmingstorm marked this pull request as ready for review September 30, 2026 19:32
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant