From 585c5ce38f86074ee5b156130807a39744d7c1fe Mon Sep 17 00:00:00 2001 From: Noisemaker111 <139656120+Noisemaker111@users.noreply.github.com> Date: Thu, 10 Sep 2026 14:28:21 -0400 Subject: [PATCH 1/2] Align command specialist purpose, training scope and handoff evidence --- AGENTS.md | 15 ++ README.md | 9 +- .../command_specialist/ENGLISH_HANDOFF.md | 10 +- experiments/command_specialist/PURPOSE.md | 128 ++++++++++++++++++ experiments/command_specialist/README.md | 13 +- .../command_specialist/codex/README.md | 5 + .../command_specialist/codex/RESULTS.md | 5 + .../codex/TEN-STAGE-PILOT.md | 5 + 8 files changed, 184 insertions(+), 6 deletions(-) create mode 100644 experiments/command_specialist/PURPOSE.md diff --git a/AGENTS.md b/AGENTS.md index 9d62399..bd55b5d 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -6,6 +6,21 @@ work, inspect real outcomes, repair ordinary failures, and return compact verifi evidence with retrievable raw output. Each delegation gets a fresh, bounded worker lifecycle. A generated program or claimed success is not completion. +Read [project purpose and evidence](experiments/command_specialist/PURPOSE.md) +before choosing the next experiment. The frontier owns reasoning and delegation; +the worker owns bounded mechanical implementation and verified evidence. The +486-example inspection/evidence adapter is not an English code-execution-trained +model. Keep product intent, trained capability and measured behavior distinct. + +Prioritize a correct grounded handoff through the real frontier host, then a +matched normal-shell comparison. Do not substitute model-size, training-speed, +or local inference experiments for that outcome. Tie supporting experiments to +an observed failure or bottleneck and a stated decision. Preserve the successful +single-handoff smoke and failed expanded workload as separate evidence. Never +infer savings from failed tasks, force baseline call counts, or treat missing +input contents as something the model should guess. Distinguish proposed, +unmerged, merged and host-verified behavior in documentation. + Work autonomously: choose the next observed failure or bottleneck, measure it, make the smallest useful improvement, verify the actual user operation and saved result, then continue. Coordinate with other agents and existing PRs instead of diff --git a/README.md b/README.md index dcd2813..0670bb0 100644 --- a/README.md +++ b/README.md @@ -61,14 +61,19 @@ These are regex heuristics checked by hand on samples. Expect a few percent nois ## Files -- `SKILL.md` — instructions the agent follows -- `reference.md` — baseline numbers and verified harness facts from the first run +- `SKILL.md` — instructions the agent follows +- `reference.md` — baseline numbers and verified harness facts from the first run - `scripts/extract.py`, `analyze.py`, `scratch.py`, `build.py`, `template.html` MIT license. ## Local command model experiment +The current [command-specialist direction](experiments/command_specialist/PURPOSE.md) +is English delegation of bounded command/cell work from a frontier agent to a +fresh local worker, returning verified evidence. It distinguishes the narrow +trained inspection pilot from the broader runtime and its measured host results. + The optional [command specialist pilot](experiments/command_specialist/README.md) audits training-data quality, benchmarks small local models on bounded PowerShell inspection and evidence-selection tasks, and trains a local LoRA adapter. Its diff --git a/experiments/command_specialist/ENGLISH_HANDOFF.md b/experiments/command_specialist/ENGLISH_HANDOFF.md index c59dc18..ee381ef 100644 --- a/experiments/command_specialist/ENGLISH_HANDOFF.md +++ b/experiments/command_specialist/ENGLISH_HANDOFF.md @@ -1,5 +1,10 @@ # English handoff contract (prototype v1) +Read [project purpose and evidence](PURPOSE.md) for the distinction between +the trained inspection adapter, the English runtime and measured frontier results. +This document describes its named prototype or historical experiment, not general +command-execution reliability. + A frontier supplies one English task, exact target bindings, relevant runtime/API context, constraints, and an observable completion condition. The frontier does not supply generated source, escaped shell strings, or a command sequence. @@ -83,8 +88,9 @@ The full contract must accommodate this without the frontier writing code: The frontier supplies checker and recorder bindings, the actual banner/cancel markers and concise node-pty API context. Completion requires observations from the real recorder and reopened saved report, not a replacement printing markers. -This Python-only prototype cannot yet perform that workload, host-native -permissions, or configured frontier return integration. The earlier file adapter +This Python-only prototype cannot yet perform that workload or provide full +host-native per-command permissions. Configured frontier return integration has +been exercised separately through the [Codex MCP prototype](codex/RESULTS.md). The earlier file adapter remains separate pending a verified replacement; its result is not proof of English delegation. diff --git a/experiments/command_specialist/PURPOSE.md b/experiments/command_specialist/PURPOSE.md new file mode 100644 index 0000000..a5dc098 --- /dev/null +++ b/experiments/command_specialist/PURPOSE.md @@ -0,0 +1,128 @@ +# Command specialist: purpose and evidence + +This is the current project direction. Historical reports describe their own +experiments; they do not redefine the product. The goal is to reduce frontier +agent time and token use on command/cell work while preserving task correctness. +Training throughput, model size and isolated inference speed are supporting +measurements, not the product outcome. + +## Division of work + +The frontier owns the user goal, higher-level reasoning, task decomposition and +whether a result is sufficient to continue. It delegates a bounded mechanical +job in English, with exact known paths, constraints, relevant environment/API +facts and a checkable result. It should not have to write the command or program +that the local worker is supposed to produce. + +The local worker owns the delegated implementation: inspect actual inputs, +construct commands or code within the supported executor, execute, inspect the +result and repair ordinary failures. One delegation creates one fresh worker +lifecycle; it can contain several internal actions. It returns compact observed +results and a retrievable raw artifact to the same frontier turn, then ends. +Warm model weights are compatible with fresh task context. A later delegation +receives explicit inputs and prior artifacts, not hidden conversation memory. + +The runtime owns exact path binding, input inspection, execution limits, process +cleanup, saved evidence and completion checks. Unknown contents, schemas and API +behavior must be inspected or reported as missing context. They are not details +for the model to guess. Verification must check the requested result: exit zero +or a printed success marker does not establish that saved artifacts are correct. +The frontier owns recovery after an unsuccessful handoff; measure that cost too. + +Python is the current trusted local execution prototype, not the definition of +the product. The original PTY startup/cancel job remains a representative future +boundary. The project is neither restricted forever to five reads nor committed +to turning the local worker into an independent general-purpose coding agent. + +## What was trained versus what was tested + +| Layer | Actual scope | What it does not establish | +| --- | --- | --- | +| Pretrained Qwen2.5-Coder model | Existing coding/instruction abilities before this project | Reliable performance on every local task | +| Preserved pilot LoRA | 486 examples: 366 typed inspection plans and 120 evidence-line selections | Post-training for arbitrary Python writing, repair or multi-step English execution | +| Original inspection runner | Five read/search/list/JSON-field operations and faithful evidence extraction | A complete command/cell delegation product | +| Codex English prototype | A real configured MCP handoff to a fresh local Python worker and return to the originating frontier | Transparent shell replacement, all native per-command permission semantics, or desktop UI verification | +| Expanded ten-stage experiment | Chained/grouped mechanical work with independent artifact assessment | Ten independently arising frontier decisions or ten required native tool calls | + +The 486 rows comprised 420 synthetic fixture examples and 66 examples with +synthetic instructions derived from historical read commands. The data pipeline +later recovered 70,611 command observations, not 70,611 gold training examples. +Recorded exit zero does not prove task correctness. See [data results](DATA_RESULTS.md). +The adapter's training objective and the English code-generation objective are +different. A failed broader task cannot by itself establish a model-size ceiling. + +## Why the successful frontier handoff was possible + +The [single-handoff CSV smoke](codex/RESULTS.md) used the real Codex host and +preserved FP16 specialist adapter. Both arms passed original and changed-input +checks. Normal shell work took 71.075 seconds; English delegation took 43.118 +seconds, including 2.813 seconds inside the local worker. Frontier native calls +fell from six to one, alongside one English handoff. + +That success demonstrates a working handoff/return path and one successful task. +The pretrained model already had coding ability; the harness could use that even +though the project adapter was trained on a narrower objective. We did not isolate +the adapter's causal contribution. Fewer frontier interactions are a plausible +source of savings, but this single pair also includes different startup, approval +and recovery costs. It does not prove typical savings or general training success. + +The [ten-stage pilot](codex/TEN-STAGE-PILOT.md) changed the workload. Normal Codex +batched it into five native calls and passed all original/alternate checks in +93.998 seconds. Ten handoffs took 200.381 seconds and two grouped handoffs took +156.575 seconds; both failed all stage assessments. One local worker satisfied a +stdout marker without creating the required artifact. That exposed a completion +contract defect, not successful execution. Dependent stage failures are not ten +independent estimates of model accuracy. + +These observations coexist: one handoff worked and was faster in its smoke; +the expanded task did not. Neither observation should be erased or generalized. +We drifted by treating broader code generation as if it had already been trained, +then pursuing prompts, model capacity and training throughput before closing the +handoff's grounding and correctness gaps. + +## Current implementation versus ongoing experiments + +The merged prototype and its recorded smokes are documented in the +[English contract](ENGLISH_HANDOFF.md) and [Codex hookup](codex/README.md). +Local work has explored explicit input inspection, artifact acceptance checks, +raw prompt changes, adapter removal, a separate 3B FP16 model and training-speed +profiling. These are diagnostics, not a new verified end-to-end result. Some +runtime changes remain unmerged; a new description is not evidence they are +available through the configured host. Preserve their raw failures and source +identities. Do not transfer old measurements to a changed harness. + +The 3B diagnostic runs but partly offloads to CPU at 32K context on the test GPU; +it is not an established replacement. A short-example checkpoint experiment sped +up training on 32 old training rows with identical saved adapter bytes. It neither +trains the English execution objective nor establishes product latency savings. +No new 5,000-example English-execution training run has been completed. + +## Next work and decision gates + +1. Re-establish one grounded English handoff through the real frontier host. + Supply observable inputs, preserve actual execution errors, check saved task + artifacts and replay on changed inputs. Reopen the result received by the + originating frontier. Start with the previous successful task shape before + expanding it. Do not count a local-only probe as this verification. +2. Compare that frozen setup with the same frontier doing normal shell work. + Let the baseline batch naturally. Repeat in alternating order, retain failures, + and include repair/fallback in total time through frontier completion. Report + frontier cached/uncached input and output separately from local model tokens; + token counts are not measured charges. Report independent evaluator time + separately from user-operation time. +3. Extend to several independently arising jobs and dependent jobs. Distinguish + mechanical stages, frontier decisions, handoff events and internal actions. + More handoffs are not inherently better. Choose boundaries for useful work, + not to inflate the baseline or projected savings. +4. If failure analysis requires training, build execution-verified examples for + the actual runtime protocol: grounded English jobs, valid implementations, + actual failures and repairs, and faithful return evidence. Keep frozen held-out + sessions/families separate; never turn the benchmark answers into training + targets. Counts alone do not establish coverage. Evaluate accuracy and whole + handoff cost before promoting a new adapter or model. + +Only investigate a lower-level optimization when an observed handoff failure or +bottleneck motivates it. State the connection and the success criterion first. +Do not automatically launch more model comparisons or training-speed work because +an earlier diagnostic made those measurements available. User steering can change +priority; update this direction explicitly instead of silently replacing the goal. diff --git a/experiments/command_specialist/README.md b/experiments/command_specialist/README.md index ffae69e..d0a2265 100644 --- a/experiments/command_specialist/README.md +++ b/experiments/command_specialist/README.md @@ -1,4 +1,12 @@ -# Local command specialist pilot +# Command specialist + +Start with [project purpose and evidence](PURPOSE.md) for the current goal, +frontier/worker responsibilities, training scope and next experiment. This page +documents the original inspection pilot. The broader [English runtime](ENGLISH_HANDOFF.md) +and [real Codex hookup](codex/README.md) are separate prototypes; their code-writing +behavior must not be described as a capability this pilot adapter was trained for. + +## Original inspection pilot This experiment tests whether a small local model can translate short inspection requests into commands and select useful evidence from output. It is a Windows @@ -23,7 +31,8 @@ The product goal is an agent-to-specialist handoff: the frontier agent supplies intent and known context, the local specialist handles supported mechanical work, and the agent receives compact faithful evidence with a retrievable raw result. These five inspection operations are a capability pilot, not the complete command -runner product or a new chat interface. Automatic host integration remains pending. +runner product or a new chat interface. An explicit Codex MCP English prototype +now exists; automatic interception of native shell work remains unimplemented. Exact paths can now be supplied separately from task wording. The runner binds them to request-local references for prediction and restores the original paths before diff --git a/experiments/command_specialist/codex/README.md b/experiments/command_specialist/codex/README.md index cbd7956..456668b 100644 --- a/experiments/command_specialist/codex/README.md +++ b/experiments/command_specialist/codex/README.md @@ -1,5 +1,10 @@ # Codex English-delegation hookup +Read [project purpose and evidence](../PURPOSE.md) for the distinction between +the trained inspection adapter, the English runtime and measured frontier results. +This document describes its named prototype or historical experiment, not general +command-execution reliability. + This connects the local FP16 worker to a **fresh Codex chat**, and captures a comparison against Codex using its normal shell/file tools. It does not compare two local models. It is a Windows, trusted-workspace Python-task prototype. diff --git a/experiments/command_specialist/codex/RESULTS.md b/experiments/command_specialist/codex/RESULTS.md index 2583000..8119c77 100644 --- a/experiments/command_specialist/codex/RESULTS.md +++ b/experiments/command_specialist/codex/RESULTS.md @@ -1,5 +1,10 @@ # Configured Codex hookup observations +Read [project purpose and evidence](../PURPOSE.md) for the distinction between +the trained inspection adapter, the English runtime and measured frontier results. +This document describes its named prototype or historical experiment, not general +command-execution reliability. + September 10, 2026. These are development hookup smokes, not a frozen benchmark. The frontier was the installed Codex CLI with its existing ChatGPT login, `gpt-6-astra`, low reasoning, normal user configuration and workspace-write with diff --git a/experiments/command_specialist/codex/TEN-STAGE-PILOT.md b/experiments/command_specialist/codex/TEN-STAGE-PILOT.md index 37fc931..86bd609 100644 --- a/experiments/command_specialist/codex/TEN-STAGE-PILOT.md +++ b/experiments/command_specialist/codex/TEN-STAGE-PILOT.md @@ -1,5 +1,10 @@ # Ten-stage incident packet pilot +Read [project purpose and evidence](../PURPOSE.md) for the distinction between +the trained inspection adapter, the English runtime and measured frontier results. +This document describes its named prototype or historical experiment, not general +command-execution reliability. + September 10, 2026. Three sequential fresh Codex CLI sessions used gpt-6-astra, low effort, existing login/configuration and normal approval review. Local workers used shell-specialist-f16 with the unchanged pilot adapter, 32768 context, 8192 From f7d0865b0a85cda314256ad16e7730eef3593347 Mon Sep 17 00:00:00 2001 From: Noisemaker111 <139656120+Noisemaker111@users.noreply.github.com> Date: Thu, 10 Sep 2026 14:29:17 -0400 Subject: [PATCH 2/2] Preserve original UTF-8 punctuation in project overview --- README.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/README.md b/README.md index 0670bb0..c2e1e08 100644 --- a/README.md +++ b/README.md @@ -61,8 +61,8 @@ These are regex heuristics checked by hand on samples. Expect a few percent nois ## Files -- `SKILL.md` — instructions the agent follows -- `reference.md` — baseline numbers and verified harness facts from the first run +- `SKILL.md` — instructions the agent follows +- `reference.md` — baseline numbers and verified harness facts from the first run - `scripts/extract.py`, `analyze.py`, `scratch.py`, `build.py`, `template.html` MIT license.