Skip to content

feat(eval): retrieval-centric capability-discovery eval harness (ir.eval) - #18

Merged
thorwhalen merged 1 commit into
masterfrom
feature/ir-eval-harness
Jun 6, 2026
Merged

thorwhalen merged 1 commit into
masterfrom
feature/ir-eval-harness

Conversation

@thorwhalen

Copy link
Copy Markdown
Member

What

Adds ir.eval — a deterministic, offline scoring layer for the question the rest of ir raises: does retrieval actually surface the right capability? This is PR 1 of the eval harness (#12): the scoring side. Case generation is deferred to a follow-up.

Design stance — retrieval-centric, reuse ef

The ir_03 research doc targets a typed @command registry (deepdiff/AST/BFCL on argument dicts). ir's corpora are documents (skills/packages/reports) joined on artifact_id, with no argument signatures — so the eval that fits is retrieval of the right artifact.

ef.evaluation already provides every retrieval metric (recall_at_k, ndcg_at_k, precision_at_k, MRR, MAP) plus the BEIR driver evaluate_retrieval (native multi-gold via Qrels) and RetrievalEvalReport. ir.eval imports these — the bridge is a ~3-line adapter (query → [artifact_id], which ef consumes as bare doc-ids).

Surface

DiscoveryCase query + multi-gold artifact_ids (empty = abstention); save_cases/load_cases JSONL with an optional corpus-version __meta__ header
as_doc_retriever / to_qrels bridge an ir corpus to ef's retriever contract
retrieval_report pure-ef path — recall@k / NDCG@k / MRR / MAP. A/B dense vs hybrid
evaluate_discovery → DiscoveryReport one pass: metrics + failure-mode taxonomy (hit_rank_1 / surfaced_low_rank / retrieval_miss) + optional score-threshold abstention proxy. Records the ranking mode and flags a silent hybrid→dense fallback when vd is absent
distractor_robustness_curve (+ _from_cases) seeded needle-in-a-haystack accuracy vs catalog size; x-axis capped at len(scope) so it never over-claims N
validate_cases drift detection — gold ids that have left the corpus (corpora are live + machine-specific)
CLI ir eval <corpus> <cases.jsonl> [--mode] [--k] (warns on drift)

Scope / non-goals

  • No LLM, no network. Dense scoring is numpy-only; lexical/hybrid use the already-declared vd dep.
  • Case generation (back-translation + name-masking), an Inspect-AI adapter, and a typed-tool/parameter scorer are intentionally deferred.

Testing & review

  • 24 hermetic tests (light embedder; disjoint-vocab corpus for exact assertions; a confusable corpus that proves the distractor curve declines). Full suite 61 passed; ruff lint + format clean.
  • The diff went through a multi-agent adversarial review (correctness · ef-contract fidelity · design · test gaps); all confirmed findings are addressed in this PR — notably the all-abstention NaN-vs-None contract, the distractor-curve x-axis capping, degenerate-trial collapse, and the reproducibility flag for vd-absent hybrid runs.

Note for maintainers

ir.eval duplicates ef.evaluation._RETRIEVAL_METRICS (a name→fn map) because that registry is private. A small improvement opportunity in ef: expose a public RETRIEVAL_METRICS registry (or get_metric_fn(name)) so ir can import it instead of mirroring it (the mirror is comment-pinned for now).

Refs #12, #1.

…val)

Add `ir.eval` — a deterministic, offline scoring layer for "does retrieval
surface the right capability?". Retrieval-centric by design: ir's corpora are
documents (skills/packages/reports) joined on artifact_id, not a typed
function-calling registry, so the eval scores retrieval of the right artifact,
not argument-dict matching.

Reuse over reinvention: metric math (recall@k / NDCG@k / MRR / MAP) and the
BEIR driver come from `ef.evaluation`; ir adapts a corpus to ef's retriever
contract (query -> [artifact_id]) and reads back a RetrievalEvalReport.

Surface:
- DiscoveryCase (+ save/load JSONL with optional corpus-version meta header)
- as_doc_retriever / to_qrels / retrieval_report (pure-ef path)
- evaluate_discovery -> DiscoveryReport (metrics + failure taxonomy +
  optional score-threshold abstention proxy; records ranking mode and flags a
  hybrid->dense fallback when vd is absent)
- distractor_robustness_curve (seeded needle-in-a-haystack; x-axis capped to
  the corpus so it never over-claims N) + distractor_curve_from_cases
- validate_cases (drift detection: gold ids that left the corpus)
- CLI: `ir eval <corpus> <cases.jsonl> [--mode] [--k]`

Scoring-harness-first: no LLM, no network (dense path is numpy-only;
lexical/hybrid use the already-declared vd dep). Case generation
(back-translation + name-masking) is deferred to a follow-up.

24 hermetic tests (light embedder, disjoint-vocab corpus for exact assertions;
a confusable corpus to prove the distractor curve actually declines) + a demo
fixture. Full suite green; ruff clean.

Refs #12, #1.
@thorwhalen
thorwhalen merged commit 052e84a into master Jun 6, 2026
12 checks passed
@thorwhalen
thorwhalen deleted the feature/ir-eval-harness branch June 6, 2026 05:42
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant