Skip to content

feat(eval): LLM-backed case generation (ir.eval_gen) - #19

Merged
thorwhalen merged 1 commit into
masterfrom
feature/ir-eval-casegen
Jun 6, 2026
Merged

thorwhalen merged 1 commit into
masterfrom
feature/ir-eval-casegen

Conversation

@thorwhalen

Copy link
Copy Markdown
Member

What

Adds ir.eval_gen — the build-time case generator that completes the eval harness (#12). It turns a corpus into DiscoveryCase data by back-translation (description → user intents; the artifact id is the free gold label), to be scored offline by ir.eval (PR #18).

This is PR 2 of the eval harness: scoring landed in #18 (ir 0.1.3); this is generation.

Design

  • Name masking (anti-leakage). The artifact name is scrubbed from the description before generation, and any intent that still reuses the name is dropped. Masking is whitespace/separator-tolerant (ci-advisor / ci advisor / ci_advisor / ciadvisor), and the output guard is bag-of-words aligned with vd's BM25 — reordered multi-word name tokens count as a leak, a single shared content word does not. This keeps the lexical leg of hybrid from trivially matching query→gold.
  • Abstention slice. A fraction of cases are "no artifact applies" (empty gold), sized to hit a target fraction.
  • Injectable LLM. query_generator / abstention_generator are callables, so all generation logic is hermetically testable with stubs; the oa-backed defaults are lazy-imported (so import ir.eval_gen stays cheap/offline).
  • corpus_signature() anchors a frozen case file to its (live, machine-specific) corpus.
  • CLI: ir eval-gen <corpus> <out.jsonl> (needs oa; scoring stays offline).

Review & testing

The diff went through a multi-agent adversarial review (correctness · masking robustness · oa-fidelity · test gaps → verify). All confirmed findings fixed, notably:

  • _parse_lines used lstrip(charset), which corrupted real queries (3D→D, -9 degrees→degrees, .env→env) — now strips only genuine list markers.
  • Multi-word masking required adjacency, so reordered name tokens slipped the guard while BM25 still matched — now whitespace-tolerant + token-set guard.
  • Name coercion for malformed name fields; deterministic max_artifacts (sorted); input validation (k>=1, abstention_frac ∈ [0,1)).

33 hermetic eval-gen tests (94 total green); ruff clean.

Real end-to-end smoke

Generated real cases from the live skills corpus and scored them — the question that motivated "eval first":

mode NDCG@10 recall@1 recall@5
dense (MiniLM) 0.845 0.625 1.000
hybrid (MiniLM+BM25+RRF) 0.694 0.375 0.875

On name-masked, paraphrastic discovery queries, dense beats hybrid — masking strips the rare identifiers the BM25 leg needs, so its signal becomes noise. (Preliminary: n=8 gold; a full generated set is the real measurement.)

Refs #12, #1.

…ness

Add `ir.eval_gen` — the build-time companion to the offline scoring harness.
Turns a corpus into DiscoveryCase eval data by back-translation (description ->
user intents; the artifact id is the free gold label), with:

- Name masking that scrubs the artifact name out of the description BEFORE
  generation, and an output guard that drops any intent reusing the name.
  Masking is whitespace/separator tolerant (ci-advisor / ci advisor / ci_advisor
  / ciadvisor) and the guard is bag-of-words aligned with vd's BM25 (reordered
  multi-word name tokens count as a leak), so the lexical leg of hybrid can't
  trivially match query->gold.
- An abstention slice (empty gold) sized to hit a target fraction.
- An injectable LLM (query_generator / abstention_generator) so all generation
  logic is hermetically testable; oa-backed defaults are lazy-imported.
- corpus_signature() to anchor a frozen case file to its (live) corpus.
- CLI: `ir eval-gen <corpus> <out.jsonl>` (needs oa; scoring stays offline).

The diff went through a multi-agent adversarial review; all confirmed findings
are addressed, notably: _parse_lines stripped meaningful leading characters
(3D / -9 degrees / .env / 24/7) via lstrip(charset) -> now strips only genuine
list markers; multi-word name masking/guard adjacency gap (above); name
coercion for malformed name fields; deterministic max_artifacts (sorted);
input validation (k>=1, abstention_frac in [0,1)).

33 hermetic eval-gen tests (94 total green); ruff clean. Refs #12, #1.
@thorwhalen
thorwhalen merged commit 3ebf888 into master Jun 6, 2026
12 checks passed
@thorwhalen
thorwhalen deleted the feature/ir-eval-casegen branch June 6, 2026 09:36
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant