Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 8 additions & 3 deletions .env.example
Original file line number Diff line number Diff line change
@@ -1,11 +1,16 @@
# This package requires no runtime environment variables for the CLI itself.
# Copy to `.env` only if you add private overrides locally (do not commit `.env`).

# Agent-test live dogfood (`bun run agent:test:live` / `agent:test:live:compare`).
# Export the key in your shell — the agent-test CLI does not auto-load `.env`.
# Direct Cursor agent tests (`bun run agent:test:direct` / `agent:test:direct:compare`).
# Export the key in your shell — agent-test does not auto-load `.env`.
# Every test execution launches a real provider agent and can incur usage.
# CURSOR_API_KEY=

# Optional model overrides for live agent / judge
# Required instead of CURSOR_API_KEY when running with `--host claude`.
# The default Cursor judge still requires CURSOR_API_KEY unless `--no-judge` is set.
# ANTHROPIC_API_KEY=

# Optional model overrides for the agent and judge
# CURSOR_AGENT_MODEL=
# CURSOR_JUDGE_MODEL=

Expand Down
2 changes: 1 addition & 1 deletion .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -15,5 +15,5 @@ src/audit/__tests__/fixtures/plugins/**/*.mjs
src/audit/__tests__/fixtures/plugins/**/*.mjs.stamp
.skeleton/catalog.md

# Raw live compare dumps (aggregate into SUMMARY.*)
# Raw direct-agent compare dumps (aggregate into SUMMARY.*)
agent-suites/evidence/runs/
10 changes: 5 additions & 5 deletions .skeleton/review-lock.json
Original file line number Diff line number Diff line change
Expand Up @@ -5,24 +5,24 @@
"reviewedAt": "2026-09-02",
"documentHash": "sha256:f9df3e875c53499eda112a7c70d2fb76c0cbbb0dfec250dc55e6d1201b04c33f",
"reviewDependencies": {
"AGENTS.md": "sha256:f3059ccf7ba993212a97be4e628e3e4c7688739036e7b0151b2005f8dcb4393b",
"AGENTS.md": "sha256:cb166286f3b8d469149488c9bf20941cffe3edb8891b1998467d600e192d7903",
"docs/developer/validation.md": "sha256:6e843854039b7a62700ccccabea9284cc00f19ce70deb3115bf77f3acc42f94a",
"src/validate/changed.ts": "sha256:c2210201db3962ca99e603a5ece323d5ffbd8a8145ee577bfa191831f977be5a"
}
},
"AGENTS.md": {
"reviewedAt": "2026-09-02",
"documentHash": "sha256:f3059ccf7ba993212a97be4e628e3e4c7688739036e7b0151b2005f8dcb4393b",
"documentHash": "sha256:cb166286f3b8d469149488c9bf20941cffe3edb8891b1998467d600e192d7903",
"reviewDependencies": {
"package.json": "sha256:33c803d5fc5280994f12cf4b3e69c7706378d111dbd3a026b092f742b6d99c04",
"package.json": "sha256:3411d4e9efaac85b784328fef349e51fe5dc47751611c9d86f7517c56da2be7a",
"src/cli.ts": "sha256:1fc7a773f0295fd209a4d3baed536c24dc7bdf3d3c079118dad6f94a2de6026a"
}
},
"README.md": {
"reviewedAt": "2026-09-02",
"documentHash": "sha256:265b48daf1265ce64be229be53c5545c43dc882585cb2e96f6858fc0410d182d",
"documentHash": "sha256:2cddd454afcfdf9acda2ff2402b0a338ea2ad10a0bc551fa27c51e24a0cb04b6",
"reviewDependencies": {
"package.json": "sha256:33c803d5fc5280994f12cf4b3e69c7706378d111dbd3a026b092f742b6d99c04",
"package.json": "sha256:3411d4e9efaac85b784328fef349e51fe5dc47751611c9d86f7517c56da2be7a",
"src/cli.ts": "sha256:1fc7a773f0295fd209a4d3baed536c24dc7bdf3d3c079118dad6f94a2de6026a"
}
},
Expand Down
2 changes: 1 addition & 1 deletion AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -59,7 +59,7 @@ bun src/cli.ts audit docs --paths=docs/a.md --fix=doc-meta --confirm-reviewed

Optional local hooks: install [pre-commit](https://pre-commit.com/) (`brew install pre-commit` or `pipx install pre-commit`), then `pre-commit install`. Customize IDE hooks from `skeleton init` are optional — not required for audit.

Behavioral A/B dogfood (live Cursor, not part of `bun run check`): [agent-suites/README.md](agent-suites/README.md) · [refs/llm-harness.md](refs/llm-harness.md).
Behavioral A/B dogfood (direct Cursor, not part of `bun run check`): [agent-suites/README.md](agent-suites/README.md) · [refs/llm-harness.md](refs/llm-harness.md).

Consumer-facing decision table and routing: [docs/developer/validation.md](docs/developer/validation.md). Common failures: [docs/developer/troubleshooting.md](docs/developer/troubleshooting.md). Day-one setup: [docs/developer/getting-started.md](docs/developer/getting-started.md).

Expand Down
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -38,7 +38,7 @@ Does an intact Skeleton contract change agent behavior — grounding on the righ

### What we did

We ran a paired live A/B self-benchmark with [`@post-print/agent-test`](https://www.npmjs.com/package/@post-print/agent-test): `skeleton-clean` vs `skeleton-messy`. The same authored prompts compare an intact fixture with a conflicting fixture.
We ran a paired direct-agent A/B self-benchmark with [`@post-print/agent-test`](https://www.npmjs.com/package/@post-print/agent-test): `skeleton-clean` vs `skeleton-messy`. The same authored prompts compare an intact fixture with a conflicting fixture.

Scenarios covered contested grounding (conflicting docs), docs-only validation routing, canonical grounding, owned-skill routing, and customize ownership. Protocol: **N=10** sequential paired compares on 2026-07-17; McNemar on paired pass/fail; median token deltas with a bootstrap CI on the mean.

Expand Down
34 changes: 17 additions & 17 deletions agent-suites/README.md
Original file line number Diff line number Diff line change
@@ -1,19 +1,19 @@
# Agent suites (skeleton behavioral benchmark)

**Source of truth for** live A/B dogfood of the Skeleton SSOT contract via `@post-print/agent-test`.
**Source of truth for** direct-agent A/B dogfood of the Skeleton SSOT contract via `@post-print/agent-test`.

<!-- doc-meta: owner=eng | last-reviewed=2026-07-17 -->

These suites measure whether a clean Skeleton structure (registry, validation lanes, customize) improves **grounding**, **validation routing**, and **token efficiency** versus a messy control tree — not portable skill conformance (that lives in [toolbox](https://github.com/csark0812/toolbox) `agent-suites/`).
These suites measure whether a clean Skeleton structure (catalog, validation lanes, customize) improves **grounding**, **validation routing**, and **token efficiency** versus a messy control tree — not portable skill conformance (that lives in [toolbox](https://github.com/csark0812/toolbox) `agent-suites/`).

Committed stats and transcript excerpts: [`evidence/`](evidence/). Protocol SSOT: [`refs/llm-harness.md`](../refs/llm-harness.md).

## Layout

```
agent-suites/
skeleton-clean/ # profile: skeleton — registry in preamble + seeded fixture docs
skeleton-messy/ # profile: shared — emptied worktree registry + conflicting docs
skeleton-clean/ # profile: skeleton — catalog routing in preamble + canonical fixture doc
skeleton-messy/ # profile: shared — no catalog routing + conflicting fixture docs
fixtures/ # source trees for seed patches
evidence/ # SUMMARY + curated transcripts (runs/ gitignored)
```
Expand All @@ -22,36 +22,36 @@ Paired scenario **names** match across clean/messy for `--compare-pairs skeleton

| Scenario | Theme |
| ----------------------------- | --------------------------------------------------- |
| `grounding: canonical topic` | Registry-first path + webhook citation |
| `grounding: conflicting docs` | SoT winner via registry |
| `grounding: canonical topic` | Catalog-first path + webhook citation |
| `grounding: conflicting docs` | SoT winner via the generated catalog |
| `routing: docs-only change` | `validate:changed` lane |
| `routing: owned skill body` | `audit:skills` lane |
| `customize: project binding` | `.skeleton/customize/` vs editing synced `SKILL.md` |

## Commands

Requires **Node ≥ 22**, exported `CURSOR_API_KEY` (see [`.env.example`](../.env.example)), and `@post-print/agent-test` ≥ 0.2.7.
Requires **Node ≥ 22**, a direct-only `@post-print/agent-test` release, and an exported provider credential (see [`.env.example`](../.env.example)). Cursor is the suite default and requires `CURSOR_API_KEY`. Claude runs use `--host claude` and require `ANTHROPIC_API_KEY`; the default judge still requires `CURSOR_API_KEY` unless `--no-judge` is set. Every execution launches a real provider agent and can incur usage.

```bash
bun run agent:test:doctor
bun run agent:test:validate
bun run agent:test:live:compare
bun run agent:test:direct:compare
```

Optional debug (staging under `$TMPDIR` by default):

```bash
bun run agent:test:live:debug -- --suite skeleton-clean
bun run agent:test:live:compare -- --debug --out-dir "$TMPDIR/skeleton-compare"
bun run agent:test:direct:debug -- --suite skeleton-clean
bun run agent:test:direct:compare -- --debug --out-dir "$TMPDIR/skeleton-compare"
```

Live is the primary signal. Suites default to `host: "replay"` so accidental non-live runs are not CI gates — always pass `--live` (the npm scripts do). Golden replay traces are deferred until rubrics stabilize.
Direct execution is the primary signal. Both suites default to `host: "cursor"`; use `--host claude` to run the same scenarios with Claude. `agent:test:validate` is the offline configuration and path check. Do not add direct execution to deterministic CI without explicit credentials, provider-usage approval, and a budget.

**Note:** Live worktrees load **preamble context from the caller checkout** (`AGENTS.md`, profile sources). Seed patches change the **worktree disk** the agent tools see. Clean vs messy therefore differs by `profile` (`skeleton` vs `shared`) plus seeded fixture/registry state. Skill seeds use worktree path `fixture-skills/` (not `.claude/skills/`) because root `.claude/` is gitignored and harness seeding must `git add` the files.
**Note:** Direct runs load **preamble context from the caller checkout** (`AGENTS.md`, profile sources). Seed patches change the **worktree disk** the agent tools see. Clean vs messy therefore differs by `profile` (`skeleton` vs `shared`) plus the seeded fixture's SSOT shape. Skill seeds use worktree path `fixture-skills/` (not `.claude/skills/`) because root `.claude/` is gitignored and harness seeding must `git add` the files.

## Success criteria (KPIs)

Score from `compare-report.json` after paired live runs. **Protocol target: N=10** independent compares before final README claims (see gates in [`evidence/SUMMARY.md`](evidence/SUMMARY.md)).
Score from `compare-report.json` after paired direct runs. **Protocol target: N=10** independent compares before final README claims (see gates in [`evidence/SUMMARY.md`](evidence/SUMMARY.md)).

| KPI | Definition | Success |
| ------------------------- | --------------------------------------------- | --------------------------------------------------------- |
Expand All @@ -64,13 +64,13 @@ Score from `compare-report.json` after paired live runs. **Protocol target: N=10

## Dogfood SOP (deposit + aggregate)

1. Export `CURSOR_API_KEY` (CLI does not load `.env`).
1. Export `CURSOR_API_KEY` for the default Cursor host, or export `ANTHROPIC_API_KEY` and pass `--host claude`. The CLI does not load `.env`; the default judge still needs `CURSOR_API_KEY` unless disabled.
2. `bun run agent:test:doctor`
3. Run a compare into a unique out dir:

```bash
OUT="$TMPDIR/skeleton-compare-run-002"
bunx agent-test --suites-dir agent-suites --live \
bunx agent-test --suites-dir agent-suites \
--compare-pairs skeleton-clean:skeleton-messy \
--fail-on=behavior --out-dir "$OUT"
mkdir -p agent-suites/evidence/runs/$(date +%Y-%m-%d)-run-002
Expand All @@ -87,8 +87,8 @@ bun run agent:evidence:excerpt -- --run-dir agent-suites/evidence/runs/<id>
```

5. Update the metric log row in [refs/llm-harness.md](../refs/llm-harness.md).
6. If seeds drift after registry edits, re-check with `git apply --check` on `**/fixtures/seeds/*.patch`.
6. If seeds drift after catalog or fixture changes, rerun `bunx agent-test --validate-seeds --suites-dir agent-suites`.

`evidence/runs/` is gitignored. Commit `SUMMARY.*`, `transcripts/`, and `samples/` after meaningful batches.

Do **not** fold live agent-test into `bun run check`.
Do **not** fold direct agent execution into `bun run check`.
4 changes: 2 additions & 2 deletions agent-suites/evidence/README.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
# Behavioral evidence

**Source of truth for** committed artifacts from the Skeleton live A/B benchmark.
**Source of truth for** committed artifacts from the Skeleton direct-agent A/B benchmark.

<!-- doc-meta: owner=eng | last-reviewed=2026-07-17 -->

Expand All @@ -18,7 +18,7 @@
`runs/` — raw per-run `compare-report.json` (+ optional suite reports). Deposit locally, then aggregate:

```bash
# after each live compare:
# after each direct compare:
mkdir -p agent-suites/evidence/runs/$(date +%Y-%m-%d)-run-NNN
cp "$TMPDIR/skeleton-compare-run-NNN/compare-report.json" \
agent-suites/evidence/runs/$(date +%Y-%m-%d)-run-NNN/
Expand Down
2 changes: 1 addition & 1 deletion agent-suites/evidence/SUMMARY.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
# Behavioral evidence summary

**Source of truth for** aggregated Skeleton A/B live compares (`skeleton-clean` vs `skeleton-messy`).
**Source of truth for** aggregated Skeleton A/B direct-agent compares (`skeleton-clean` vs `skeleton-messy`).

<!-- doc-meta: owner=eng | last-reviewed=2026-07-17 -->

Expand Down
2 changes: 1 addition & 1 deletion agent-suites/evidence/charts/pass-rates.svg
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
2 changes: 1 addition & 1 deletion agent-suites/fixtures/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ Readable trees used to build `skeleton-clean` / `skeleton-messy` seed patches un

| Tree | Role |
| ---- | ---- |
| `clean/docs/fixture/` | Registered Billing API SoT + legacy decoy |
| `clean/docs/fixture/` | Canonical Billing API SSOT + legacy decoy |
| `clean/skill-trees/` | Bodies seeded into worktree `fixture-skills/` |
| `messy/docs/fixture/` | Conflicting SoT claims (no correct webhook) |
| `messy/skill-trees/` | Same skill bodies for messy arm |
Expand Down
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
# Billing API (legacy notes)

This draft is **not** registered. Wrong webhook: `https://legacy.example.com/hooks/bill`
This draft is **not** canonical. Wrong webhook: `https://legacy.example.com/hooks/bill`

Agents must not cite this URL when the registry points at billing-api.md.
Agents must not cite this URL when the catalog points at billing-api.md.
14 changes: 2 additions & 12 deletions agent-suites/skeleton-clean/fixtures/seeds/grounding.patch
Original file line number Diff line number Diff line change
Expand Up @@ -20,17 +20,7 @@ new file mode 100644
@@ -0,0 +1,5 @@
+# Billing API (legacy notes)
+
+This draft is **not** registered. Wrong webhook: `https://legacy.example.com/hooks/bill`
+This draft is **not** canonical. Wrong webhook: `https://legacy.example.com/hooks/bill`
+
+Agents must not cite this URL when the registry points at billing-api.md.

--- a/.skeleton/registry.md
+++ b/.skeleton/registry.md
@@ -22,6 +22,7 @@
| Plugins | [plugins.md](../docs/developer/plugins.md) |
| Customize | [customize.md](../docs/developer/customize.md) |
| Skeleton skill | [SKILL.md](../skeleton/SKILL.md) |
+| Billing API | [billing-api.md](../docs/fixture/billing-api.md) |

## Customizations
+Agents must not cite this URL when the catalog points at billing-api.md.

4 changes: 2 additions & 2 deletions agent-suites/skeleton-clean/scenarios.json
Original file line number Diff line number Diff line change
@@ -1,8 +1,8 @@
{
"name": "skeleton-clean",
"description": "Live A/B clean arm: profile=skeleton + catalog/SSOT fixture docs. Pair with skeleton-messy via --compare-pairs.",
"description": "Direct-agent A/B clean arm: profile=skeleton + catalog/SSOT fixture docs. Pair with skeleton-messy via --compare-pairs.",
"defaults": {
"host": "replay",
"host": "cursor",
"profile": "skeleton",
"skills": "none"
},
Expand Down
38 changes: 0 additions & 38 deletions agent-suites/skeleton-messy/fixtures/seeds/customize-demo.patch
Original file line number Diff line number Diff line change
Expand Up @@ -10,41 +10,3 @@ new file mode 100644
+<!-- doc-meta: owner=eng | last-reviewed=2026-07-17 -->
+
+Protocol: answer with DEMO_SKILL_BASELINE when invoked with no customize overlay.

--- a/.skeleton/registry.md
+++ b/.skeleton/registry.md
@@ -1,30 +1,10 @@
# Registry

-<!-- doc-meta: owner=eng | last-reviewed=2026-07-14 -->
+<!-- doc-meta: owner=eng | last-reviewed=2026-07-17 -->

-**Source of truth for** topic routing in this repo. Edit rows here; edit content in canonical files only.
+**Source of truth for** topic routing in this repo.

## Documentation

-| Topic | Canonical file |
-| --------------------- | ---------------------------------------------------------- |
-| Package overview | [README.md](../README.md) |
-| Agent cold-start | [AGENTS.md](../AGENTS.md) |
-| Authoring conventions | [authoring.md](../docs/authoring.md) |
-| Ecosystem tiers | [tiers.md](../docs/tiers.md) |
-| Getting started | [getting-started.md](../docs/developer/getting-started.md) |
-| Install | [install.md](../docs/developer/install.md) |
-| Config | [config.md](../docs/developer/config.md) |
-| Validation | [validation.md](../docs/developer/validation.md) |
-| Troubleshooting | [troubleshooting.md](../docs/developer/troubleshooting.md) |
-| Doc system | [doc-system.md](../docs/developer/doc-system.md) |
-| Audit | [audit.md](../docs/developer/audit.md) |
-| Plugins | [plugins.md](../docs/developer/plugins.md) |
-| Customize | [customize.md](../docs/developer/customize.md) |
-| Skeleton skill | [SKILL.md](../skeleton/SKILL.md) |
-
-## Customizations
-
-| Topic | Canonical file |
-| ------------------------------------------------------------------------------------------------------ | ------------------------------------------ |
-| Customize: skeleton-specific code-review overlays (validation ladder, invariant matrices, Action bar). | [code-review.md](customize/code-review.md) |
+| Topic | Canonical file |
+| ----- | -------------- |
Loading