Skip to content

bench: evidence matrix, router comparisons, and corpus-size influence - #191

Open
Lucas-Bur wants to merge 16 commits into
mainfrom
bench/retrieval-evidence-matrix
Open

Lucas-Bur wants to merge 16 commits into
mainfrom
bench/retrieval-evidence-matrix

Conversation

@Lucas-Bur

Copy link
Copy Markdown
Owner

Implements the #166 evidence-chain issues #160, #170, #173, and #175 as benchmark-only commits. No production retrieval behavior changes.

#160 — Application corpora (8e2e838, 6fc9de5)

Adds three real-application corpora so every parsed language now has one library and one application:

Corpus Language Kind Revision
T3 Code TypeScript application badae6a5cc8325dcd5a145bea6f7b8ac692818a1
beets Python application b7952299941543d4507ac7931edb223acd684b3d (v2.13.1)
Alacritty Rust application 94e7c8874e526b1e67b349d9ba30ddf81669119e (v0.17.0)

Each manifest carries 15 authored Pix-style queries with four query forms and exact file + symbol gold. vp run bench:retrieval:corpus validates all 6 repositories and all 90 intents against the pinned checkouts.

#170 — Matrix execution and artifact merge (c662b4a)

  • benchmarks/matrix/full.json defines the complete matrix as explicit run axes; it expands to 12,960 coordinates across 6 repositories, 3 models, 5 optimization profiles, 4 objectives, all fusion methods, grouped folds, and repository holdouts.
  • benchmarks/retrieval/matrix.ts expands the manifest deterministically into 30 process invocations and merges artifacts against an explicit plan, rejecting duplicate, missing, and unexpected coordinates. Each merged coordinate keeps the full router result, diagnostics, source timestamp, and timings.
  • Schema 32 adds artifactSerializationDurationMs, measured from the actual preflight serialization pass.
  • vp run bench:retrieval:matrix executes all 30 model-backed runs and writes the merged matrix artifact (release evidence; not part of the dev loop).

#173 — Router model and fitting-method comparison (7e4ebd4)

  • routeWithComparisonModel compares the production multiplicative router against a bounded regularized log-linear gate over the same parameter space.
  • Four derivative-free fitting methods: staged, alternating block-coordinate, deterministic Halton restarts, and local-search baseline.
  • Dimension diagnostics classify every coordinate as active, inactive, or data-constant from development evidence; inactive coordinates can be pruned.
  • Complexity-aware utility and one-standard-error simplicity selection are both recorded for every search.
  • Selection is development-only; evaluateRouterComparisonHoldout scores already-selected candidates on an excluded holdout with no reselection path.
  • Every result carries the explicit no-global-optimality claim.

#175 — Corpus-size influence (3463d81)

  • Deterministic sub-samples at 200/500/1000/2500/5000/full chunks; gold chunks of evaluated questions always stay resolvable, unavoidable drops are recorded with reasons.
  • Sub-sample corpora rebuild dense BM25/identifier indexes with remapped chunk indices; the execution test validates gold resolution on the real pinned fd corpus (vp run bench:retrieval:corpus-size).
  • The sweep records per-size optima with exact (corpus size, fusion, profile, objective, strategy, fold) coordinates and fits a per-channel log-linear model over log(chunkCount).
  • The sensitivity check compares adjacent-size weight shifts against the combined one-standard-error noise band and only recommends promoting a corpus-size factor when a shift leaves the band.

Validation

  • vp check clean (format, lint, types).
  • vp test: 71 files, 498 passed, 1 skipped.
  • vp run lint:fallow at the pre-existing finding baseline (no new dead exports or duplication).
  • vp run bench:retrieval:corpus and bench:retrieval:corpus-size green on the pinned checkouts.

@coderabbitai

coderabbitai Bot commented Aug 28, 2026 •

Copy link
Copy Markdown

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: 392cee98-a91f-47ff-8b4f-eff9eb6c7e2a


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@codecov

codecov Bot commented Aug 28, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant