Skip to content

Proposal: standalone ranx-openeval-adapter (Qrels/Run/per-query metric arrays <-> EvalPort Suite/ResultSet) #82

Description

@adhabnr-ux

Hi — I maintain EvalPort, an open, framework-agnostic JSON-Schema spec for portable LLM/IR eval data (TestCase/Suite/Grader/ResultSet documents plus a validator), so a dataset or a graded run can move between tools without a bespoke converter each time.

I read ranx/data_structures/qrels.py, ranx/data_structures/report.py, and ranx/metrics/recall.py (and skimmed the other metric files) before filing this, not just the README.

What I found genuinely maps well:

  • Qrels.qrels ({q_id: {doc_id: relevance_score}}) and Run (the analogous {q_id: {doc_id: score}} shape) are exactly the two things an EvalPort TestCase needs for a retrieval grader: one TestCase per q_id, with the qrels doc/relevance map going into retrieval_context (or metadata — EvalPort doesn't have a first-class "ground-truth ranked list" field, more below) and the run's ranked list as actual_output.
  • The per-metric low-level functions (e.g. recall._recall_parallel, and the equivalent in ndcg.py, average_precision.py, reciprocal_rank.py, etc.) return a np.ndarray of one score per query before evaluate()/compare() reduce it to the single number shown in Report. That per-query array is the part that maps cleanly onto EvalPort's GraderResult.score — the aggregate-only numbers in Report.to_dict() (results[model][metric]) would lose exactly the granularity EvalPort is meant to preserve.
  • Report.to_dict()'s comparisons (per-metric p-values between model pairs) and win_tie_loss don't have a natural home in a single ResultSet — that's model-vs-model, not test-case-vs-grader. I'd put that in ResultSet.metadata as an openeval.*-namespaced extension rather than force it into summary, since inventing a new top-level construct for pairwise stats isn't something I'd want to do without buy-in from someone who actually uses compare().

What I'm proposing, no strong preference between the two:

  1. A standalone ranx-openeval-adapter package (depends on ranx normally, ships to_openeval()/from_openeval(), validates against EvalPort's actual JSON Schema in tests) living in EvalPort's own adapters/ directory. Zero footprint on this repo.
  2. Or a PR here if you'd rather it live closer to ranx.io/ranx.meta.

Given ranx is qrels/run-centric rather than test-case-centric, I'd scope v1 to exactly what I described above (one TestCase per query, per-query metric arrays as GraderResults) and leave compare()'s statistical-significance output as an optional metadata extension rather than a forced fit. Let me know if that scoping looks right, or if this isn't useful for your roadmap — no worries either way.

— Sahi, independent contributor (not affiliated with ranx)

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions