Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
32 changes: 32 additions & 0 deletions .github/workflows/agent-checks.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,32 @@
name: Agent checks
on:
pull_request:
branches: [agents, main]
types: [opened, synchronize, reopened, ready_for_review, edited]
push:
branches: [agents]
permissions:
contents: read
concurrency:
group: agent-checks-${{ github.event.pull_request.number || github.ref }}
cancel-in-progress: true
jobs:
core:
name: Windows core
runs-on: windows-latest
timeout-minutes: 10
steps:
- uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
with:
persist-credentials: false
- uses: actions/setup-python@a26af69be951a213d495a4c3e4e4022e16d87065
with:
python-version: '3.12'
- name: Validate syntax
run: python -m compileall -q scripts experiments/command_specialist
- name: Path binding and saved evidence invariants
run: python -m unittest discover -s experiments/command_specialist -p test_bindings.py -v
- name: Native and PowerShell contract invariants
run: python -m unittest discover -s experiments/command_specialist -p test_contract.py -v
- name: Frozen main candidate and patch notes
run: python scripts/check_release.py
76 changes: 76 additions & 0 deletions AGENTS.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,76 @@
# Agents and main

`agents` is the development integration branch. `main` is the stable installation
branch. Jon authorized agents-branch automation on September 10, 2026: agents may
prepare owned changes, verify them, open ready PRs against `agents`, and merge them
when the gates below pass. This replaces per-PR human merge requests for `agents`.
Only Jon personally merges release PRs into `main`. Agents must never merge main,
enable its auto-merge, or push directly to it, even after chat approval. Publication,
deployments, model/weight changes, and production data changes need separate approval.
This agents automation applies to every agent acting for a maintainer with live
repository write, maintain, or admin permission; no particular model/host is privileged.

Start from fetched `agents` in an owned worktree. Preserve other sessions and reuse
existing PRs for the same work. Run commands from the repository root with Python
3.12. Required checks are:

```
python -m compileall -q scripts experiments/command_specialist
python -m unittest discover -s experiments/command_specialist -p test_bindings.py -v
python -m unittest discover -s experiments/command_specialist -p test_contract.py -v
```

Also verify the affected user operation through the public CLI and reopen its saved
artifact. Hosted CI has no private model, GPU, or transcripts; it supplements local
product verification. Do not run live model benchmarks concurrently or change the
shared Ollama service. Retain `shell-specialist-pilot` and its existing adapter.

`.github/workflows/agent-checks.yml` runs the required `Windows core` check on PRs
into both branches and pushes to `agents`. It runs on a hosted Windows runner with
read-only permissions and no persisted checkout credentials. PR code never receives
a privileged merge token. `agents` requires this GitHub Actions check and an up-to-date PR, prevents force
pushes/deletion, and applies protection to administrators. No approving review is
required for `agents`. Main requires the same CI/PR gate; Jon performs the merge
himself. A separate approving reviewer is not required because Jon authors these
PRs through the same account. The platform cannot distinguish Jon from an agent
using his credentials; the human-only merge rule is enforced by this workflow and
the coordinator's unconditional refusal of main.

The local coordinator is `scripts/integrate_agents.py`. After reviewing the diff,
verifying the current head locally, and waiting for CI, run:

```
python scripts/integrate_agents.py --pr <number> --verified-head <verified-head>
```

It accepts only ready same-repository PRs targeting `agents` when both the PR author
and the acting account have live write/maintain/admin permission; checks the exact head, workflow identity/result, and mergeability; and
merges with GitHub's expected-head condition. It refuses `main`. `--check-only`
checks eligibility without merging. Invoke it from a trusted checkout; passing a
head is the agent's attestation of actual local verification, not proof supplied
by CI. Every maintainer's agent uses this same automatic integration path; it does not
require Jon to run the command. It is agent-operated, not an unattended scheduler.
The agent continues through merge and merged-revision verification without asking
Jon for routine agents integration approval. Each successful merge automatically
runs `scripts/prepare_release.py`: it generates patch notes from merged change
titles and opens a frozen main release PR. An existing open release batch is
preserved; later agents changes wait for the next batch. Release preparation may
be resumed with `python scripts/prepare_release.py`; this never merges main.

This repository is a local skill/CLI, with no hosted frontend, backend, database,
queue, production service, or deployment trigger configured here. Dev validation
runs the merged `agents` revision from an isolated worktree with fresh fixtures and
ignored `work/` artifacts. Stable users continue to install `main`; existing local
installations and sessions are not repointed. No automatic host integration is
claimed. Source revision, loaded CLI source, and saved artifact must agree when
reporting dev verification.

For stable promotion, branch a frozen release candidate from a verified agents
revision and open a ready PR into `main`. Include the full diff, user-facing notes,
verification evidence, and rollback target (the prior main revision). Later agents
changes must not join that candidate silently. CI checks the candidate identity
and generated patch notes via `scripts/check_release.py`, in addition to core
checks. Ask Jon to merge the CI-green release PR himself; agents stop before main. The agents GitHub PR/check gates are enforced, but this account is shared
by Jon and agents: human authorization is a workflow rule, not an independently
verified GitHub reviewer identity. The coordinator cannot merge stable. No package
or runtime deployment is implied by either merge.
34 changes: 34 additions & 0 deletions experiments/command_specialist/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -45,6 +45,40 @@ A caller can hand off a UTF-8 JSON task file:
python experiments/command_specialist/run.py --root '<directory>' --task-file task.json --backend native
```

Runtime limits are explicit CLI options and keyword arguments on `inspect_request`.
Defaults remain `--num-ctx 4096 --num-predict 160` for compatibility. To evaluate
more context/output capacity with the same model and adapter:

```powershell
python experiments/command_specialist/run.py --root '<directory>' --task-file task.json --backend native --num-ctx 8192 --num-predict 2048
```

These values reach **both** planning and evidence requests; Modelfile defaults do
not override them. 8192/2048 is an evaluation profile, not a universal capacity
recommendation or a speed claim. `--evidence-max-lines` (100),
`--evidence-max-chars` (8000), and `--packet-max-chars` (2000) independently bound
selection input and compact stdout/stderr. All limits must be positive integers.
The character window includes newline separators. It is not a tokenizer budget:
callers must allow room for their intent, numbered evidence, system instructions,
and generated output within the chosen context. Ollama may truncate overlong
prompts; this pilot does not prove arbitrary inputs fit from character counts.

A line-limited selection is marked truncated; a character-overflow window skips
selection and asks the caller to inspect the raw artifact. Increasing model context
alone does not enlarge the evidence window. `fallback: "inspect_raw_result"`
identifies bounded or rejected output; it does not trigger an automatic retry or
host action. Raw stdout/stderr remain complete for the executed bounded operation
(the plan's requested line count is still part of that operation).

Empty stdout deterministically returns empty evidence without a second model call.
For nonempty evidence, the artifact records the requested selection, model response,
selection timing, and returned packet. Planning output stopped at the generation
limit is rejected before execution and saved with its timing and raw response;
evidence output stopped at the limit falls back to the already saved raw result.
The artifact records effective runtime limits, execution time, prompt/decode token
counts and durations, and Ollama's stop reason. No model or service settings change.
See [runtime handoff evaluation](RUNTIME_RESULTS.md) for measured scope and limits.

Or use the Python `run.inspect_request(root, request, targets=...)` boundary. For a
small CLI request, `--target 'build=logs/build.log' --request 'Read the last 5 lines
of {{build}}.'` supplies the same binding. Existing unbound `--request` calls retain
Expand Down
134 changes: 134 additions & 0 deletions experiments/command_specialist/RUNTIME_RESULTS.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,134 @@
# Public handoff runtime evaluation

Measured September 10, 2026. The unchanged `shell-specialist-pilot` model is
Qwen2.5-Coder 1.5B Q4_K_M with its existing LoRA adapter. Current checks confirmed
Ollama 0.33.3, RTX 3070 8GB, Ryzen 9 3900X, and GPU execution. Shared Ollama service
configuration, weights, quantization, and adapter were not changed. No deployment
or host integration occurred.

## Outcome

The confirmed bug was unnecessary model selection on empty stdout. In the merged
handoff, both no-match cases generated 160 tokens of nonexistent line numbers,
stopped at the output limit, and fell back with invalid JSON. Deterministic empty
selection fixes those cases and removes the second inference call. Nonempty
evidence still uses the model and is returned verbatim.

The public CLI and callable now expose context/output and evidence/packet window
limits. Defaults remain compatible at 4096/160; 8192/2048 is an opt-in evaluation
profile. Raw artifacts now retain evidence selection/timing, effective settings,
execution timing, stop reasons, and raw model responses. Output-limit planning
failures save diagnostic artifacts and execute nothing. Selection failure keeps
the pre-selection raw artifact. Character-window accounting includes newlines.

## Final paired measurements

Three runs used the same saved 16-case synthetic set across all treatments,
with fresh opaque filenames and two registered candidates per case. All five
operations were included, along with longer caller context, evidence selection,
two empty searches, a character-window overflow, compact-output truncation, and
an escaping target rejected before inference. PowerShell reference outputs were
prepared outside measurement. JSON correctness compares values; other operations
compare actual stdout and exit status. Evidence compares requested source lines
and verbatim text; explicit overflow fallback and invalid-input rejection have
separate expected outcomes. Exact plan text is not the correctness definition.

Each treatment block ran a full-workload warmup and then the same measured cases.
Order was merged/compat/capacity, capacity/compat/merged, then merged/compat/capacity.
Blocks avoid forcing context reload on every case. These are paired by saved case,
with alternating treatment order, not randomized interleaving or independent samples.

| Run | Source/profile | Checks | CLI median ms | CLI p95 ms | Mean generated tokens | Mean prompt ms | Mean decode ms | Weighted decode tokens/s |
|---|---|---:|---:|---:|---:|---:|---:|---:|
| 1 | Merged PR2, 4096/160 | 14/16 | 289.6 | 1186.1 | 40.94 | 37.6 | 226.8 | 180.53 |
| 1 | Revised, 4096/160 | 16/16 | 278.6 | 450.3 | 20.94 | 29.9 | 116.1 | 180.31 |
| 1 | Revised, 8192/2048 | 16/16 | 292.7 | 418.8 | 20.94 | 30.4 | 119.3 | 175.57 |
| 2 | Merged PR2, 4096/160 | 14/16 | 289.4 | 1166.3 | 40.94 | 33.1 | 221.8 | 184.58 |
| 2 | Revised, 4096/160 | 16/16 | 287.4 | 438.1 | 20.94 | 29.2 | 119.1 | 175.78 |
| 2 | Revised, 8192/2048 | 16/16 | 283.9 | 419.2 | 20.94 | 28.5 | 117.7 | 177.94 |
| 3 | Merged PR2, 4096/160 | 14/16 | 295.7 | 1191.8 | 40.94 | 36.1 | 225.3 | 181.67 |
| 3 | Revised, 4096/160 | 16/16 | 298.6 | 463.3 | 20.94 | 32.7 | 125.4 | 166.96 |
| 3 | Revised, 8192/2048 | 16/16 | 296.7 | 482.3 | 20.94 | 29.5 | 120.5 | 173.76 |

Whole-handoff time includes Python/CLI startup, task-file reading, binding, model
calls, validation, native execution, real raw-artifact writes, packet output, and
artifact reload. The loopback diagnostic proxy buffers requests/replies in memory;
trace serialization, report writes, and correctness scoring occur after the timer.
Proxy overhead remains included equally. Prompt/decode figures sum both model calls
per handoff, including rejected generations. Means include the pre-inference invalid
case as zero tokens/time. Weighted throughput is total tokens / total decode seconds,
not the arithmetic mean of per-request throughput. No decode improvement is claimed.

Across 48 measured attempts per treatment, execution was correct on all 45 executable
cases and all three invalid cases were rejected. Evidence/fallback checks passed
42/48 for merged code and 48/48 for each revised profile. The six merged failures
were empty-search evidence selection, not wrong execution or filename resolution.
There were no observed execution/evidence regressions. Nonempty evidence selection
passed throughout this set. Each revised arm had six flagged truncations/fallbacks
(three compact-output windows, three oversized selection windows); merged had 12,
including its six selection failures. Fallback means the caller must inspect the
raw artifact; no automated retry or host fallback is installed.

Pooled medians/p95 were 291.2/1186.1 ms merged, 288.0/438.1 ms revised compatibility,
and 290.1/418.8 ms revised capacity. Median differences are small and inconsistent
across runs. The lower tail latency and generation count come from removing the
failed empty-selection calls. With only 16 cases per run, nearest-rank p95 is the
maximum observation; do not generalize it to production. No >5% decode-throughput
win, general reliability claim, or whole-agent speed claim is supported.

## Cache conditions and capacity checks

The initial unchanged public-CLI smoke returned the correct execution and selected
error/summary lines; its raw artifact was reopened. It took 2379.4 ms including
2001.4 ms model load. This is a cold-load observation, not a warm baseline.

The final paired table reports warmed repeated workloads. First traversal blocks
were retained separately: median 299.3 ms merged, 284.0 ms compatibility, and
299.2 ms capacity; largest load times were 2866.9, 5.9, and 2891.4 ms respectively.
These are not controlled cold-cache comparisons: context changes can reload the
runner, later arms reuse prompts, and fresh filenames become short references.
No shared-service unload/cache reset was performed. Novel real-world prompt and
cold-start performance remain unestablished.

Separate final-source public-CLI probes (not pooled into the paired table):

- 100 error lines: the 160-token selection cap was reached and flagged; raw stdout
survived. At 2048, all 100 source lines were selected verbatim using 296 tokens.
- Longer caller context: Ollama reported 5187 prompt tokens at context 8192; actual
execution and requested error/summary evidence were correct.
- A two-line evidence window omitted the third-line summary with `truncated: true`
and raw-result fallback. A five-character evidence window skipped selection.
- A five-character packet window returned the prefix, flagged truncation, and kept
complete raw stdout. A one-token planning cap exited nonzero, saved the raw model
reply/stop reason, and executed nothing. Zero context was rejected before inference.
- The final public CLI also passed through the actual PowerShell backend, with
readable bound paths, source lines 2/3, and the saved output reopened.

These probes demonstrate particular capacities, not that arbitrary prompts fit.
Character windows are not token counts; the pilot does not detect all server-side
prompt truncation. The 100-line operation contract remains. Larger context alone
does not enlarge selection windows. Unexpected provider/network error recovery and
arbitrary Unicode/JSON equivalence beyond the existing pilot remain unproven.

## Evidence and reproduction

Local evidence remains under this task worktree in `work/runtime-final/`:
`evaluate.py`, immutable `source-before/` and `source-after/`, `manifest.json`,
saved tasks/fixtures and PowerShell references, all 288 warmup/measured records,
per-operation raw artifacts and buffered wire traces, `summary.json`, and nine
CLI probes under `probes/`. SHA-256 source/input checks passed after measurement;
the four executed production modules match the final worktree byte-for-byte.
The exploratory set remains separately in `work/runtime-evaluation/`. The final
set was generated after implementation was frozen and not used to tune the model
or code. Historical development/final-test corpora were never opened.

To reproduce locally, copy the saved evaluator and both source snapshots into a
new empty result directory, then run `python <new-directory>/evaluate.py` from the
repository. It creates new filenames, freezes inputs, and performs all three runs.
The evaluator is local diagnostic evidence, not a retained product scenario suite.
Raw inputs, model traces, corpora, and weights are excluded from Git.

Validation: all 10 focused binding/contract invariant tests passed, including actual
native and PowerShell execution, path confinement, revalidation, literal quoting,
configured settings on both model calls, and persisted faithful evidence. This does
not validate automatic OpenCode2 integration; that remains a subsequent bounded step.
19 changes: 16 additions & 3 deletions experiments/command_specialist/benchmark.py
Original file line number Diff line number Diff line change
Expand Up @@ -22,17 +22,30 @@ def post(base: str, path: str, body: dict) -> dict:
return json.load(response)


def predict(base: str, model: str, case: dict) -> tuple[dict, dict]:
class PredictionError(ValueError):
def __init__(self, message, timing):
super().__init__(message)
self.timing = timing


def predict(base: str, model: str, case: dict, *, num_ctx=4096, num_predict=160) -> tuple[dict, dict]:
start = time.perf_counter()
response = post(base, "/api/chat", {
"model": model, "messages": messages(case), "stream": False,
"format": prediction_schema(case),
"think": False, "keep_alive": "5m",
"options": {"temperature": 0, "seed": 20260909, "num_ctx": 4096, "num_predict": 160},
"options": {"temperature": 0, "seed": 20260909, "num_ctx": num_ctx, "num_predict": num_predict},
})
timing = {"inference_wall_ms": (time.perf_counter() - start) * 1000}
timing.update({key: response.get(key) for key in ["total_duration", "load_duration", "prompt_eval_count", "prompt_eval_duration", "eval_count", "eval_duration"]})
return json.loads(response["message"]["content"]), timing
timing["done_reason"] = response.get("done_reason")
timing["raw_response"] = response["message"]["content"]
if response.get("done_reason") == "length":
raise PredictionError("Model output limit reached; no prediction accepted", timing)
try:
return json.loads(response["message"]["content"]), timing
except (ValueError, TypeError) as error:
raise PredictionError(str(error), timing) from error


def equivalent_output(case: dict, actual: str, expected: str) -> bool:
Expand Down
Loading
Loading