When qrels and/or run are passed to ranx.evaluate() as plain Python dicts (a typed and documented input form, Dict[str, Dict[str, Number]]), each is converted to a positional list and every metric kernel pairs qrels[i] with run[i] by index. check_keys() only checks that the two arguments contain the same set of query ids, never that they are in the same order. The sort=True argument of python_dict_to_typed_list() only sorts each query's own document array, never the outer per-query list. If the two dicts list their queries in different insertion order, every per-query score and the resulting mean are silently computed against another query's relevance judgments. This is routine when they come from two separately loaded JSON files, two DataFrame groupbys, or two dict comprehensions over differently ordered sources.
The same happens when a Qrels/Run object is mixed with a raw dict. The object side is key-sorted by sort_dict_by_key before conversion, but the dict side keeps its insertion order. check_keys still passes, because numba's typed Dict.keys() is a collections.abc.KeysView and compares to dict_keys as a set.
ranx/utils.py:24-42 (docstring omitted, list construction condensed)
def python_dict_to_typed_list(x: Dict[str, Dict[str, Number]], sort: bool = True):
out = TypedList([
np.array([[hash(doc_id), score] for doc_id, score in doc.items()], dtype=np.float64)
for doc in x.values() # outer order = the dict's insertion order
])
if sort:
out = descending_sort_parallel(out) # sorts each query's own doc array only
return out
ranx/meta/evaluate.py:59-62 (the only cross-check performed)
def check_keys(qrels, run):
assert (
qrels.keys() == run.keys()
), "Qrels and Run query ids do not match. ..." # keys-view equality is a set comparison
ranx/meta/evaluate.py:134-138 and :146
if isinstance(qrels, (Qrels, dict)) and isinstance(run, (Run, dict)):
check_keys(qrels, run)
_qrels = convert_qrels(qrels) # dict -> python_dict_to_typed_list(qrels, sort=True); Qrels -> key-sorted
_run = convert_run(run) # dict -> python_dict_to_typed_list(run, sort=True); Run -> key-sorted
...
metric_scores_dict[metric] = metric_switch(m)(_qrels, _run, k, rel_lvl) # line 146: pairs _qrels[i] with _run[i]
Measured
| Case |
qrels order |
run order |
check_keys |
Reported mean precision@1 |
Correct mean precision@1 |
| Control A (same order) |
q_bad, q_good |
q_bad, q_good |
passes |
1.0 |
1.0 |
| Bug (run dict order swapped) |
q_bad, q_good |
q_good, q_bad |
passes |
0.0 |
1.0 |
Bug (key-sorted qrels, as from Qrels(), + raw dict run) |
a, b |
b, a |
passes |
0.0 |
1.0 |
Control B (both key-sorted, as Qrels()/Run() do) |
q_bad, q_good |
q_bad, q_good |
passes |
1.0 |
1.0 |
In both bug rows every query retrieves exactly its own single relevant document, so the run is perfect, yet evaluate() reports the worst possible score, 0.0, with no exception or warning. These numbers come from python_dict_to_typed_list, descending_sort, check_keys, clean_qrels, fix_k, _hits, _precision and sort_dict_by_key copied verbatim, with numba decorators removed, numba.typed.List replaced by list, and the typed-dict sort replaced by an equivalent plain-dict key sort. They ran in a plain Python process with only numpy.
Consequence
Any evaluate() call where at least one argument is a raw dict silently returns scores keyed to the wrong query whenever the dict's key order differs from the other argument's order: the other dict's insertion order, or the key-sorted order of a Qrels/Run object. test_python_dict and test_python_dict_2 pass raw dicts, but both build qrels and run with the same query order. test_python_dict_2 shuffles document order inside each query, but not query order. test_std goes through Qrels(dict)/Run(dict). Nothing covers differing query order. Only calls where both arguments are Qrels/Run objects are immune, because both sides are key-sorted before conversion. Typed-list and ndarray inputs carry no query ids and are positional by definition.
Suggested fix
Apply the canonicalization the Qrels/Run path already performs to the raw-dict branch: in convert_qrels/convert_run, or in python_dict_to_typed_list, iterate the dict in sorted-key order. That way a dict's insertion order can never decide which query's ground truth a score is computed against, and dict inputs line up with key-sorted Qrels/Run objects in the mixed case. Alternatively, make check_keys order-sensitive (list(qrels.keys()) == list(run.keys())) so a mismatch raises the existing AssertionError instead of mis-scoring. The first option also fixes the mixed object/dict case without extra user action.
Happy to open the PR.
When
qrelsand/orrunare passed toranx.evaluate()as plain Python dicts (a typed and documented input form,Dict[str, Dict[str, Number]]), each is converted to a positional list and every metric kernel pairsqrels[i]withrun[i]by index.check_keys()only checks that the two arguments contain the same set of query ids, never that they are in the same order. Thesort=Trueargument ofpython_dict_to_typed_list()only sorts each query's own document array, never the outer per-query list. If the two dicts list their queries in different insertion order, every per-query score and the resulting mean are silently computed against another query's relevance judgments. This is routine when they come from two separately loaded JSON files, two DataFrame groupbys, or two dict comprehensions over differently ordered sources.The same happens when a
Qrels/Runobject is mixed with a raw dict. The object side is key-sorted bysort_dict_by_keybefore conversion, but the dict side keeps its insertion order.check_keysstill passes, because numba's typedDict.keys()is acollections.abc.KeysViewand compares todict_keysas a set.ranx/utils.py:24-42(docstring omitted, list construction condensed)ranx/meta/evaluate.py:59-62(the only cross-check performed)ranx/meta/evaluate.py:134-138and:146Measured
check_keysQrels(), + raw dict run)Qrels()/Run()do)In both bug rows every query retrieves exactly its own single relevant document, so the run is perfect, yet
evaluate()reports the worst possible score, 0.0, with no exception or warning. These numbers come frompython_dict_to_typed_list,descending_sort,check_keys,clean_qrels,fix_k,_hits,_precisionandsort_dict_by_keycopied verbatim, with numba decorators removed,numba.typed.Listreplaced bylist, and the typed-dict sort replaced by an equivalent plain-dict key sort. They ran in a plain Python process with onlynumpy.Consequence
Any
evaluate()call where at least one argument is a raw dict silently returns scores keyed to the wrong query whenever the dict's key order differs from the other argument's order: the other dict's insertion order, or the key-sorted order of aQrels/Runobject.test_python_dictandtest_python_dict_2pass raw dicts, but both build qrels and run with the same query order.test_python_dict_2shuffles document order inside each query, but not query order.test_stdgoes throughQrels(dict)/Run(dict). Nothing covers differing query order. Only calls where both arguments areQrels/Runobjects are immune, because both sides are key-sorted before conversion. Typed-list and ndarray inputs carry no query ids and are positional by definition.Suggested fix
Apply the canonicalization the
Qrels/Runpath already performs to the raw-dict branch: inconvert_qrels/convert_run, or inpython_dict_to_typed_list, iterate the dict in sorted-key order. That way a dict's insertion order can never decide which query's ground truth a score is computed against, and dict inputs line up with key-sortedQrels/Runobjects in the mixed case. Alternatively, makecheck_keysorder-sensitive (list(qrels.keys()) == list(run.keys())) so a mismatch raises the existingAssertionErrorinstead of mis-scoring. The first option also fixes the mixed object/dict case without extra user action.Happy to open the PR.