Skip to content

Commit bc55001

Browse files
anandgupta42claude
andauthored
test: learn benchmark harness, task project and lesson pools (#1407)
* test: [#1406] learn benchmark harness, task project and lesson pools Moved out of #1405 unchanged (from `feat/rsi-workspace-learning`), so the product PR stays reviewable. Produces the numbers in `research/rsi-workspace-learning-2026-09-30/learn-v1-results.md`. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test: [#1406] harden the learn benchmark harness Review fixes for research tooling; the benchmark definition (task project, seed data, `to_utc` macro and verifier checks) is unchanged so the published runs stay comparable. - safety: validate ids and paths before any deletion; fake backend binds to 127.0.0.1, debug routes need a token, authorization gaps closed - reproducibility: no machine-specific paths (dbt, CLI, worktrees); env overrides for the CLI and models; the fake backend is the default and the real SaaS needs an explicit opt-in and workspace id - scoring: runs whose agent turn did not complete no longer count; topic sessions must complete both turns in one session; analysis keeps unique directory identities - README with prerequisites (needs the `learn` features from #1405), suites and env vars; self-tests for the harness, v1bench and fake backend Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test: [#1406] address the second round of harness review comments Benchmark definition (verifier checks, task project, seeds, macros) unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test: [#1406] restore the published benchmark configuration; third review round Model, arm, lesson and pool defaults match the commit that produced the published runs (780168a); comparison checkouts are required env vars. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test: [#1406] require an explicit workspace id for real-SaaS harness runs Removes the internal workspace id default from `run_corrections.sh` and `run_drift.sh`; SaaS mode already refuses to run without a positive `WORKSPACE_ID`, and fake-backend runs do not use it. Self-test fixtures use a neutral id. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test: [#1406] combine topic-switch leak scans; safer workdir reset and baseline default - topic-switch records combine first- and second-turn leak scans - prepare_workdir deletes only a destination it created (marker + git origin) - run_corrections evaluates a fresh baseline unless one is given explicitly Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test: [#1406] fix five harness review findings - publish_replace requires an explicit workspace ID (flag or WORKSPACE_ID) - bootstrap keeps training sessions from iterations 10 and later - ablation retries start from fresh learning state - topic-switch rejects unknown session IDs before touching output - leaked records make an arm incomplete and are excluded from analysis Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test: [#1406] fix two regressions from the previous harness fix - ablation validates --from-loop before resetting learning state - publish_replace requires a workspace ID only for the SaaS backend Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test: [#1406] reset ablation state only after preflight; scope WORKSPACE_ID to SaaS - ablation keeps prior learning state when task, user or model setup fails - publish_replace reads WORKSPACE_ID only for the SaaS backend Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
1 parent cc22cda commit bc55001

138 files changed

Lines changed: 14673 additions & 0 deletions

File tree

Some content is hidden

Large Commits have some content hidden by default. Use the searchbox below for content that may be hidden.
Lines changed: 91 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,91 @@
1+
# RSI research harness
2+
3+
Research tooling only. Model-driven suites incur provider charges; the self-tests below do not.
4+
The learning loops, v1 delivery, bootstrap, and drift suites require the **`learn` features from
5+
PR #1405**. This harness-only branch does not supply them. Use a checkout containing those
6+
features through `ALTIMATE_CMD`; changing product code is outside this experiment.
7+
The verifier, fake backend, and local `none`/`gold`/`playbook` evaluation arms do not need `learn`.
8+
9+
Prerequisites: Bun, Git, repository dependencies installed for the selected CLI, and Python
10+
3.10+ for the dbt environment (harness syntax supports 3.9+). The local checks used Python 3.11.
11+
For example, from the repository root:
12+
13+
```bash
14+
python3.11 -m venv /tmp/rsi-dbt
15+
/tmp/rsi-dbt/bin/pip install dbt-core==1.12.5 dbt-duckdb==1.11.0
16+
export DBT_BIN=/tmp/rsi-dbt/bin/dbt
17+
export DBT_PYTHON=/tmp/rsi-dbt/bin/python
18+
```
19+
20+
Pin package versions and record the CLI commit and model IDs alongside any reported results.
21+
The shared default CLI is this repository's `packages/opencode/src/index.ts`; baseline and fix
22+
comparison wrappers retain their historical `rsi-base` and `rsi` worktrees. An override is a
23+
shell-quoted command string, e.g. `ALTIMATE_CMD="bun run --conditions=browser '/path with spaces/src/index.ts'"`.
24+
Model-driven runs need provider authentication usable from the isolated homes. Account/org policies
25+
may block Claude, including Claude on Vertex; choose models your account can use. The harness checks
26+
requested model IDs before evaluation and treats timeouts, errors and incomplete turns as failed runs.
27+
28+
## Running suites
29+
30+
Free local checks (from the repository root):
31+
32+
```bash
33+
experiments/rsi-workspace/demo/run_task.sh selftest
34+
python3 experiments/rsi-workspace/fake-backend/selftest.py
35+
python3 experiments/rsi-workspace/harness/selftest.py
36+
python3 experiments/rsi-workspace/harness/v1bench/selftest.py
37+
```
38+
39+
Model-driven suites, from `experiments/rsi-workspace/harness/`:
40+
41+
```bash
42+
bash run_all.sh my-ci-loop # train / validation gate / publish / heldout arms
43+
bash run_corrections.sh my-corrections # simulated teammate corrections / publish / eval
44+
python3 eval.py --arm gold --runs 1 --run-id my-eval --out runs/my-eval/eval/gold.jsonl
45+
bash run_arms.sh my-ci-loop # rerun final arms of an existing completed loop
46+
python3 publish_replace.py runs/my-ci-loop # recover publishing using supported publish/update
47+
bash budget/run_budget.sh # rule-budget arms
48+
bash budget/run_budget2.sh # scope/conflict arms
49+
bash budget/run_drift.sh my-drift "${REFLECTOR_MODEL:-google-vertex/gemini-3.5-flash}"
50+
FIX_SRC_ROOT=/path/to/fix-checkout bash budget/run_drift_matrix.sh
51+
FIX_SRC_ROOT=/path/to/fix-checkout bash budget/run_drift_matrix2.sh
52+
bash v1bench/run_baselines.sh # n50-short / all-50 / all-300 / all-1000
53+
bash v1bench/run_all_v1.sh # retrieval, vague task, topic, bootstrap, drift suites
54+
bash v1bench/run_fix_v1.sh # focused fix comparison (ALTIMATE_CMD selects checkout)
55+
```
56+
57+
See [v1 suite plan](harness/v1bench/PLAN.md), [topic switch](harness/v1bench/topic_switch/README.md),
58+
[vague tasks](harness/v1bench/vague_tasks/README.md), [demo](demo/README.md), and
59+
[fake backend](fake-backend/README.md) for task definitions and additional switches.
60+
`fake-backend/demo.sh` runs the model-driven two-user workspace demo.
61+
IDs must be simple names (letters/digits plus `.`, `_`, `-`); empty, absolute and traversing IDs
62+
are rejected. Main loops require a fresh run ID. Recovery/evaluation replaces complete arm output;
63+
v1 resumes only outputs containing the exact unique expected task/run keys.
64+
65+
## Environment
66+
67+
| Variable | Default / meaning |
68+
|---|---|
69+
| `ALTIMATE_CMD` | Bun CLI in this repository; honors quoted paths |
70+
| `DBT_BIN`, `DBT_PYTHON` | `dbt` on PATH; Python beside dbt, or explicitly selected interpreter |
71+
| `AGENT_MODEL` | Shared Python drivers: `google-vertex-anthropic/claude-haiku-4-5@20251001`; v1 and budget drift shell drivers: `google-vertex/gemini-3.5-flash` |
72+
| `REFLECTOR_MODEL`, `REVIEWER_MODEL` | Shared default: `google-vertex-anthropic/claude-sonnet-4-6@default`; drift shell reviewer: `google-vertex/gemini-3.1-pro-preview`; independently overridable |
73+
| `BACKEND` | Historical default `saas`, requiring explicit opt-in; `fake` uses loopback and isolated test users |
74+
| `ALLOW_REAL_SAAS` | Unset; must equal `1` to permit real workspace access |
75+
| `WORKSPACE_ID` / `--workspace-id` | Required positive ID for SaaS; binding must resolve to it |
76+
| `SAAS_CREDS_DIR` | No default; explicit directory containing `home-a/.altimate/altimate.json` and `home-b/...` for SaaS |
77+
| `RUNS`, `PARALLEL`, `K`, `RUNS_VAL` | Wrapper sample count, concurrency, iterations and validation repetitions |
78+
| `AGENT_TIMEOUT`, `STAGGER_SECONDS` | Agent deadline (600s), startup spacing (3s) |
79+
| `BASELINE_FROM` | Corrections driver defaults to `runs/saas-v2/eval/none.jsonl`; set empty to evaluate anew; retain original logs when reusing |
80+
| `FIX_SRC_ROOT`, `STRONG_MODEL`, `WEAK_MODEL` | Drift matrix checkout (historical `rsi-fix` worktree by default), reflector choices |
81+
| `BUDGET_REAL_FILE` | Optional real-lesson source; defaults to `runs/corr-main/playbooks/final.md` |
82+
| `FAKE_DEBUG_TOKEN` | Standalone fake-server admin token; unset disables debug routes; harness generates its own |
83+
84+
SaaS publishing can update an existing **owned** `team-playbook`. Only run it against an intended test
85+
workspace, with `BACKEND=saas ALLOW_REAL_SAAS=1 WORKSPACE_ID=... SAAS_CREDS_DIR=...` (Python entrypoints use
86+
`--backend saas --workspace-id ...`). No real SaaS or paid benchmark is needed for self-tests.
87+
88+
Lesson pools/playbooks are deterministic: regenerate with `python3 harness/v1bench/pool.py` then
89+
`python3 harness/v1bench/make_playbooks.py` from this directory. Committed pools, playbooks, lesson IDs/order, and benchmark treatment defaults match `780168ab18`.
90+
Historical result Markdown predates these verifier and delivery fixes; it is retained as historical evidence,
91+
not rescored or claimed as current benchmark results.
Lines changed: 133 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,133 @@
1+
# acme-shop RSI demo: does the agent learn team conventions over time?
2+
3+
For humans running the experiment. **The agent never sees `demo/`**; it only gets a workdir made by
4+
`prepare_workdir.py` (a copy of `project/`) and the one-line ticket prompt.
5+
6+
```
7+
demo/
8+
project/ dbt-duckdb project "acme_shop" (the agent's repo; README does not document the conventions)
9+
verifier/
10+
check.py hidden CI: check.py <workdir> <task_id> -> JSON
11+
tasks/*.json 14 tasks: id, split, prompt, source, table, target_model, check_type, setup, verify
12+
gold/ hand-written gold solutions (train-refunds, heldout-invoices, 2 controls)
13+
gold_playbook.md upper-bound arm: the conventions as a SKILL.md body
14+
selftest.py gold must pass; naive / over-applied / garbage must fail
15+
prepare_workdir.py <task_id> <dest>: copy project, git init (+fake origin), commit, dbt seed. Idempotent.
16+
run_task.sh prepare | check | selftest | list wrapper
17+
```
18+
19+
## Commands
20+
21+
```bash
22+
python3.11 -m venv /tmp/rsi-dbtenv
23+
/tmp/rsi-dbtenv/bin/pip install dbt-core==1.12.5 dbt-duckdb==1.11.0
24+
export DBT_BIN=/tmp/rsi-dbtenv/bin/dbt # default: dbt on PATH
25+
# DBT_PYTHON defaults to the python next to DBT_BIN (needs duckdb), else the current interpreter
26+
./run_task.sh list
27+
./run_task.sh prepare train-refunds /tmp/wd # fresh workdir; prints the ticket prompt. Re-running rebuilds it.
28+
# ... agent works in /tmp/wd with the prompt ...
29+
./run_task.sh check train-refunds /tmp/wd # JSON; exit 0 iff all checks ok (~6s, never modifies /tmp/wd)
30+
./run_task.sh selftest # ~55s; --all-gold solves all 12 staging tasks programmatically; --all-naive
31+
```
32+
33+
A generic harness reads each `tasks/<id>.json`: `setup` and `verify` are argv templates (run from `demo/`, `{workdir}`
34+
substituted). Output: `{"task_id","pass","score","checks":[{"name","kind","ok","message"}]}`; `score` = fraction of checks ok,
35+
`kind` is `lint` (message may name the rule) or `data` (symptom only). Per-check `ok` gives per-convention scores.
36+
37+
The verifier copies the workdir to a temp dir, replaces `seeds/`, overlays pristine `macros/`, restores
38+
`dbt_project.yml` and `profiles.yml`, then reseeds a fresh duckdb and builds there. Seed edits, edits to tracked
39+
macros, and a stale db are discarded; workdir-only macros survive the overlay and can affect the build.
40+
41+
## Hidden conventions (the ground truth)
42+
43+
| id | convention | check kind |
44+
|----|-----------|-----------|
45+
| C1 | `models/staging/<source>/stg_<source>__<entity>.sql` (entity = plural table minus `raw_`) | lint |
46+
| C2 | PK `id` -> `<singular>_id`; model declared in `_<source>__models.yml` with `unique` + `not_null` on it | lint |
47+
| C3 | `*_cents` columns via `{{ cents_to_dollars() }}`, renamed without `_cents`; no `*_cents` in output (values verified) | lint |
48+
| C4 | every timestamp column via `{{ to_utc() }}`, output named `<stem>_at` (`refunded_ts`->`refunded_at`, `created`->`created_at`); `date` columns exempt | data |
49+
| C5 | sources with `_is_deleted`: `where not _is_deleted` and column not exposed (verified by row count in duckdb) | data |
50+
| C6 | `dbt build --select <model>` passes (model + tests) | lint |
51+
52+
Controls (conventions must NOT be over-applied): `control-customers-vip` (add boolean `is_vip` = >=3 orders to
53+
`stg_shop__customers`; K1 build+tests, K2 values, K3 existing columns/rows unchanged, K4 `stg_shop__orders` untouched) and
54+
`control-payments-by-month` (ad-hoc `analyses/payments_by_month.sql`, totals must stay in cents; K1 exists, K2 compiles,
55+
K3 columns, K4 numbers match). Applying `cents_to_dollars` or `_is_deleted` filtering blindly fails them.
56+
57+
## Splits
58+
59+
| split | task | source | C3 money | C4 timestamps | C5 soft delete |
60+
|-------|------|--------|:--:|:--:|:--:|
61+
| train | train-refunds | shop | x | x | x |
62+
| train | train-payments | billing | x | x | |
63+
| train | train-shipments | shop | x | x | x |
64+
| train | train-coupons | shop | x | | x |
65+
| val | val-subscriptions | billing | x | x | x |
66+
| val | val-products | shop | x | x | x |
67+
| val | val-plans | billing | x | x | |
68+
| val | val-csat-surveys | support | | x | x |
69+
| heldout | heldout-invoices | billing | x | x | x |
70+
| heldout | heldout-support-tickets | support | | x | x |
71+
| heldout | heldout-disputes | billing | x | x | x |
72+
| heldout | heldout-ledger-entries | billing | x | x | |
73+
| control | control-customers-vip | shop | existing model; no convention applies | | |
74+
| control | control-payments-by-month | billing | analysis; keep cents, no convention applies | | |
75+
76+
Every learning split exercises C3, C4 and C5. No non-staging "heldout" variant was added (it could not be made fair:
77+
the conventions are staging-specific); the two control tasks play that role.
78+
79+
Note on `heldout-support-tickets`: the first runs showed every arm (including the gold-playbook arm) naming the model
80+
`stg_support__tickets` instead of the expected `stg_support__support_tickets`, which failed all of C1-C6 and scored 0.00
81+
everywhere. The prompt now names the entity ("...; the entity is support_tickets") to remove that ambiguity. C1 remains
82+
strict on the name; when the expected file is missing and exactly one new/changed `stg_support__*.sql` exists (vs the
83+
pristine `project/`), C2-C6 are evaluated on that model so the task still measures the conventions. Zero or several
84+
candidates keep the all-fail behaviour.
85+
86+
## Fair-evaluation notes
87+
88+
- **Inferable from existing code** (`stg_shop__customers/orders`, `_shop__models.yml`, `macros/`): C1 file naming/location,
89+
the CTE pattern, C2 (`customer_id`/`order_id` PK rename, YAML location, `unique`+`not_null`), and C3 (orders uses
90+
`cents_to_dollars` and drops `_cents`). Also `macros/to_utc.sql` exists, so the macro is discoverable, but no model uses it.
91+
- **Not inferable from the repo**: C4 (use of `to_utc`, the `_at` suffix) and C5 (soft-delete filtering; no existing source
92+
has `_is_deleted`). These can only be learned from verifier feedback (or the playbook). C4/C5 feedback is deliberately
93+
symptom-only (e.g. "returned 40 rows; reconciliation expects 37"); C1-C3 feedback names the rule, like a linter.
94+
- First attempts should therefore pass C1/C2/C3/C6 often and fail C4/C5; gains over time should show on C4/C5 in val/heldout.
95+
- C3/C4 text checks are regexes on comment-stripped SQL (`cents_to_dollars(col)`, `to_utc(col)`); equivalent hand-written SQL fails them.
96+
C4 expects `<col minus _ts>_at`; any other `_at` spelling is reported as violating the suffix convention.
97+
- Timestamp columns are detected from the seed types (`timestamp`); `date` columns (`order_date`, `expires_on`) are exempt.
98+
- `train-coupons` has no timestamp column and `*-plans`, `ledger-entries`, `payments` have no `_is_deleted`; the checks report "not applicable" (ok).
99+
- Over-eager agents that blanket-apply `where not _is_deleted` to customers fail the control build; those converting to dollars in the analysis fail K4.
100+
- Self-test results below were produced with dbt-core 1.12 / dbt-duckdb; each verifier call takes about 6s.
101+
102+
## Self-test results (`./run_task.sh selftest`)
103+
104+
```
105+
gold:train-refunds pass=True score=1.0 (C1..C6 all ok; 22 rows == 25 minus 3 deleted)
106+
gold:heldout-invoices pass=True score=1.0 (37 rows == 40 minus 3 deleted)
107+
gold:control-customers-vip pass=True score=1.0 (5 VIPs)
108+
gold:control-payments-by-month pass=True score=1.0 (35 month/method rows match)
109+
naive:train-refunds pass=False score=0.333 (C1, C6 ok; C2..C5 FAIL)
110+
naive:heldout-invoices pass=False score=0.333
111+
overapplied:control-payments-by-month pass=False score=0.75
112+
overapplied:control-customers-vip pass=False score=0.25
113+
empty/garbage workdir, unknown task pass=False (no crash; every check reports a message)
114+
./run_task.sh selftest --all-gold: all 12 staging tasks solved by a generator pass (score 1.0)
115+
SELFTEST OK
116+
```
117+
118+
Naive solution (`select * from {{ source('shop', 'refunds') }}`, no YAML) on `train-refunds`:
119+
120+
```
121+
[ok] C1_location_naming (lint): found models/staging/shop/stg_shop__refunds.sql
122+
[FAIL] C2_primary_key_and_tests (lint): stg_shop__refunds: model is not declared in models/staging/shop/_shop__models.yml. Team rule: each source folder has one _shop__models.yml with its models; column `refund_id` is missing unique and not_null test(s) in models/staging/shop/_shop__models.yml; primary key `id` is not renamed to `refund_id` (Team rule: <singular_entity>_id).
123+
[FAIL] C3_money_cents_to_dollars (lint): stg_shop__refunds: column `amount_cents` exposed raw. Team rule: money columns must be converted with {{ cents_to_dollars() }} and renamed without the _cents suffix (amount); column(s) `amount_cents` still carry the _cents suffix. Team rule: no *_cents columns in staging output.
124+
[FAIL] C4_timestamps_utc_at (data): stg_shop__refunds: column `refunded_ts` is not timezone-normalized; the downstream join against the finance calendar (UTC) fails on it; timestamp column(s) `refunded_ts` pass through un-normalized and un-renamed.
125+
[FAIL] C5_soft_deletes (data): stg_shop__refunds: returned 25 rows; the reconciliation against the source system expects 22. 3 row(s) should not reach analytics; exposes internal column `_is_deleted`, which must not reach analytics.
126+
[ok] C6_dbt_build (lint): `dbt build --select stg_shop__refunds` succeeded (model and its tests)
127+
```
128+
129+
Over-applied control (`cents_to_dollars` in the analysis):
130+
131+
```
132+
[FAIL] K4_numbers_match_finance_export (data): 70 month/method row(s) differ from the finance export, e.g. ('2025-01-01', 'bank_transfer'): got payments/total_cents ('1', '94') vs expected ('1', '9410')
133+
```

0 commit comments

Comments
 (0)