feat(eval): LLM-backed case generation (ir.eval_gen) - #19
Merged
Merged
Conversation
…ness Add `ir.eval_gen` — the build-time companion to the offline scoring harness. Turns a corpus into DiscoveryCase eval data by back-translation (description -> user intents; the artifact id is the free gold label), with: - Name masking that scrubs the artifact name out of the description BEFORE generation, and an output guard that drops any intent reusing the name. Masking is whitespace/separator tolerant (ci-advisor / ci advisor / ci_advisor / ciadvisor) and the guard is bag-of-words aligned with vd's BM25 (reordered multi-word name tokens count as a leak), so the lexical leg of hybrid can't trivially match query->gold. - An abstention slice (empty gold) sized to hit a target fraction. - An injectable LLM (query_generator / abstention_generator) so all generation logic is hermetically testable; oa-backed defaults are lazy-imported. - corpus_signature() to anchor a frozen case file to its (live) corpus. - CLI: `ir eval-gen <corpus> <out.jsonl>` (needs oa; scoring stays offline). The diff went through a multi-agent adversarial review; all confirmed findings are addressed, notably: _parse_lines stripped meaningful leading characters (3D / -9 degrees / .env / 24/7) via lstrip(charset) -> now strips only genuine list markers; multi-word name masking/guard adjacency gap (above); name coercion for malformed name fields; deterministic max_artifacts (sorted); input validation (k>=1, abstention_frac in [0,1)). 33 hermetic eval-gen tests (94 total green); ruff clean. Refs #12, #1.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Adds
ir.eval_gen— the build-time case generator that completes the eval harness (#12). It turns a corpus intoDiscoveryCasedata by back-translation (description → user intents; the artifact id is the free gold label), to be scored offline byir.eval(PR #18).This is PR 2 of the eval harness: scoring landed in #18 (ir 0.1.3); this is generation.
Design
ci-advisor/ci advisor/ci_advisor/ciadvisor), and the output guard is bag-of-words aligned withvd's BM25 — reordered multi-word name tokens count as a leak, a single shared content word does not. This keeps the lexical leg of hybrid from trivially matching query→gold.query_generator/abstention_generatorare callables, so all generation logic is hermetically testable with stubs; theoa-backed defaults are lazy-imported (soimport ir.eval_genstays cheap/offline).corpus_signature()anchors a frozen case file to its (live, machine-specific) corpus.ir eval-gen <corpus> <out.jsonl>(needsoa; scoring stays offline).Review & testing
The diff went through a multi-agent adversarial review (correctness · masking robustness ·
oa-fidelity · test gaps → verify). All confirmed findings fixed, notably:_parse_linesusedlstrip(charset), which corrupted real queries (3D→D,-9 degrees→degrees,.env→env) — now strips only genuine list markers.namefields; deterministicmax_artifacts(sorted); input validation (k>=1,abstention_frac ∈ [0,1)).33 hermetic eval-gen tests (94 total green); ruff clean.
Real end-to-end smoke
Generated real cases from the live skills corpus and scored them — the question that motivated "eval first":
On name-masked, paraphrastic discovery queries, dense beats hybrid — masking strips the rare identifiers the BM25 leg needs, so its signal becomes noise. (Preliminary: n=8 gold; a full generated set is the real measurement.)
Refs #12, #1.