Hi — I maintain EvalPort, an open, framework-agnostic JSON-Schema spec for portable LLM/IR eval data (TestCase/Suite/Grader/ResultSet documents plus a validator), so a dataset or a graded run can move between tools without a bespoke converter each time.
I read ranx/data_structures/qrels.py, ranx/data_structures/report.py, and ranx/metrics/recall.py (and skimmed the other metric files) before filing this, not just the README.
What I found genuinely maps well:
Qrels.qrels ({q_id: {doc_id: relevance_score}}) and Run (the analogous {q_id: {doc_id: score}} shape) are exactly the two things an EvalPort TestCase needs for a retrieval grader: one TestCase per q_id, with the qrels doc/relevance map going into retrieval_context (or metadata — EvalPort doesn't have a first-class "ground-truth ranked list" field, more below) and the run's ranked list as actual_output.
- The per-metric low-level functions (e.g.
recall._recall_parallel, and the equivalent in ndcg.py, average_precision.py, reciprocal_rank.py, etc.) return a np.ndarray of one score per query before evaluate()/compare() reduce it to the single number shown in Report. That per-query array is the part that maps cleanly onto EvalPort's GraderResult.score — the aggregate-only numbers in Report.to_dict() (results[model][metric]) would lose exactly the granularity EvalPort is meant to preserve.
Report.to_dict()'s comparisons (per-metric p-values between model pairs) and win_tie_loss don't have a natural home in a single ResultSet — that's model-vs-model, not test-case-vs-grader. I'd put that in ResultSet.metadata as an openeval.*-namespaced extension rather than force it into summary, since inventing a new top-level construct for pairwise stats isn't something I'd want to do without buy-in from someone who actually uses compare().
What I'm proposing, no strong preference between the two:
- A standalone
ranx-openeval-adapter package (depends on ranx normally, ships to_openeval()/from_openeval(), validates against EvalPort's actual JSON Schema in tests) living in EvalPort's own adapters/ directory. Zero footprint on this repo.
- Or a PR here if you'd rather it live closer to
ranx.io/ranx.meta.
Given ranx is qrels/run-centric rather than test-case-centric, I'd scope v1 to exactly what I described above (one TestCase per query, per-query metric arrays as GraderResults) and leave compare()'s statistical-significance output as an optional metadata extension rather than a forced fit. Let me know if that scoping looks right, or if this isn't useful for your roadmap — no worries either way.
— Sahi, independent contributor (not affiliated with ranx)
Hi — I maintain EvalPort, an open, framework-agnostic JSON-Schema spec for portable LLM/IR eval data (
TestCase/Suite/Grader/ResultSetdocuments plus a validator), so a dataset or a graded run can move between tools without a bespoke converter each time.I read
ranx/data_structures/qrels.py,ranx/data_structures/report.py, andranx/metrics/recall.py(and skimmed the other metric files) before filing this, not just the README.What I found genuinely maps well:
Qrels.qrels({q_id: {doc_id: relevance_score}}) andRun(the analogous{q_id: {doc_id: score}}shape) are exactly the two things an EvalPortTestCaseneeds for a retrieval grader: oneTestCaseperq_id, with the qrels doc/relevance map going intoretrieval_context(ormetadata— EvalPort doesn't have a first-class "ground-truth ranked list" field, more below) and the run's ranked list asactual_output.recall._recall_parallel, and the equivalent inndcg.py,average_precision.py,reciprocal_rank.py, etc.) return anp.ndarrayof one score per query beforeevaluate()/compare()reduce it to the single number shown inReport. That per-query array is the part that maps cleanly onto EvalPort'sGraderResult.score— the aggregate-only numbers inReport.to_dict()(results[model][metric]) would lose exactly the granularity EvalPort is meant to preserve.Report.to_dict()'scomparisons(per-metric p-values between model pairs) andwin_tie_lossdon't have a natural home in a singleResultSet— that's model-vs-model, not test-case-vs-grader. I'd put that inResultSet.metadataas anopeneval.*-namespaced extension rather than force it intosummary, since inventing a new top-level construct for pairwise stats isn't something I'd want to do without buy-in from someone who actually usescompare().What I'm proposing, no strong preference between the two:
ranx-openeval-adapterpackage (depends onranxnormally, shipsto_openeval()/from_openeval(), validates against EvalPort's actual JSON Schema in tests) living in EvalPort's ownadapters/directory. Zero footprint on this repo.ranx.io/ranx.meta.Given
ranxis qrels/run-centric rather than test-case-centric, I'd scope v1 to exactly what I described above (oneTestCaseper query, per-query metric arrays asGraderResults) and leavecompare()'s statistical-significance output as an optional metadata extension rather than a forced fit. Let me know if that scoping looks right, or if this isn't useful for your roadmap — no worries either way.— Sahi, independent contributor (not affiliated with ranx)