You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
bpref returns NaN for qrels with no judged non-relevant docs (trec_eval returns a defined value), and fisher_randomization_test reports p_value=0.0 / significant=True whenever a per-query score is NaN #85
When a query's qrels contain no judged non-relevant documents (n_non_rels == 0), min(n_rels, n_non_rels) is 0. Every retrieved relevant document then contributes 0/0, and the query's score is nan. Positive-only qrels are a common shape: MS MARCO-style qrels files, for example, list only relevant pairs. Under @njit no warning reaches the caller.
This was reported in #66 and closed as expected behaviour ("with no known irrelevant documents you are dividing by zero … N is zero"). The maintainer also said ranx's bpref "should be the same as TREC Eval". trec_eval does not divide by zero here. m_bpref.c special-cases it explicitly:
/* Special case nonrel_so_far == 0 to avoid division by 0 */if (nonrel_so_far>0) {
bpref+=1.0-
(((double) MIN(nonrel_so_far, res_rels.num_rel)) /
(double) MIN(num_nonrel, res_rels.num_rel));
} elsebpref+=1.0;
...
if (res_rels.num_rel)
bpref /= res_rels.num_rel;
With num_nonrel == 0, nonrel_so_far is always 0. Each retrieved relevant document adds 1.0, and trec_eval returns (relevant retrieved) / R. On identical input, ranx returns nan. This matches the #66 reporter's observation that trec_eval and pytrec_eval produced values for the same data.
The NaN then hits a separate problem in ranx/statistical_tests/fisher_randomization_test.py:29-46:
If any per-query score is nan, both control_treatment_diff and every permuted_diff are nan, and nan >= nan is False. counter_array therefore never increments, and the test returns p_value = 0.0, significant = True: the strongest possible "the systems differ" result, whatever the data. The bug is independent of bpref. Any NaN in the score arrays triggers it.
Measured
Each case re-runs the function bodies copied verbatim from HEAD 7363db0c (v0.3.21) under numba 0.67 / numpy 2.5.3 / scipy 1.18.1, with @njit / parallel=True kept. Doc IDs are small floats, as in ranx's float64 qrels/run arrays.
#
Case
Result
Expected
1
_bpref on tests/unit/ranx/metrics/bpref_test.py q1 (qrels {d2:1,d5:1,d1:0,d4:0}, run d1,d2,d3,d4) - control
0.25
0.25 (matches the test's assertion)
2
np.isclose(0.25, 0.999) - must-fail control
False
False
3
qrels [[2,1],[5,1],[7,1]] (relevant only), run retrieves 2, 5, 7 first (perfect ranking) plus one unjudged doc
nan
1.0 (trec_eval: 3 relevant retrieved / R=3)
4
Same qrels, run = one unjudged doc, then doc 2
nan
0.333… (trec_eval: 1/3)
5
Same qrels, run retrieves no relevant doc
0.0
0.0 (no division reached)
6
Public path _bpref_parallel over [q1 fixture, case-3 query]
[0.25, nan], no warning
[0.25, 1.0]
7
Public path, query with only non-relevant judgments (n_rels == 0) - control
0.0
0.0 (trec_eval also returns 0)
8
np.mean([nan, 0.5, 0.8, 0.9]) (the aggregation used by evaluate() and compare())
nan
a per-query NaN should not silently poison the run mean
an error or a NaN-flagged result, not maximal significance
10
Control: same arrays with the nan replaced by 0.51
p_value=0.705, significant=False
not significant
11
Control: [0.0]*10 vs [1.0]*10, 2000 permutations
p_value=0.0005, significant=True
significant
12
paired_student_t_test path (ttest_rel then p_value <= max_p, line 15) on the case-9 arrays
p_value=nan, significant=False
distinguishable from "not significant"
Consequence
One query whose qrels list only relevant documents makes that query's bprefnan. evaluate() then reports the run-level bpref as nan (ranx/meta/evaluate.py:151, :162, :167, plain np.mean), and so does compare() (ranx/meta/compare.py:93, float(np.mean(...))), with nothing pointing at the query or the cause. trec_eval returns a defined score for the same input.
compare(..., stat_test="fisher") passes the raw per-query arrays (compare.py:96) to fisher_randomization_test. For every pair of runs whose scores for that metric include a nan, it reports p_value: 0.0, significant: True, which makes two indistinguishable systems (case 9 vs. control case 10) look maximally significantly different. The same happens for any NaN-producing metric, not only bpref. (compare() defaults to stat_test="student", so this path needs stat_test="fisher"; the internal compute_statistical_significance defaults to "fisher".)
On the default student path the NaN becomes significant: False. The return value cannot distinguish that from a genuine non-significant result.
Suggested fix
_bpref: match trec_eval. A relevant document with no judged non-relevant document ranked above it contributes 1.0, so the 0/0 never happens. For example, compute the per-relevant term only where cumsum(hit_list == 0) > 0 and use 1.0 otherwise, or return (relevant retrieved) / n_rels when n_non_rels == 0. Add an explicit if n_rels == 0: return 0.0 in the same style as the sibling guards (hits.py:13-15, recall.py:14-16, f1.py:14-16, average_precision.py:13-15, r_precision.py:16-17). Standalone @njit_bpref currently raises ZeroDivisionError in that case and relies on the parallel wrapper to yield 0.0. If the current NaN behaviour is kept on purpose, a clear error or warning naming the query (as suggested in I'm getting NaN for the BPref measurement #66) would at least make it visible.
fisher_randomization_test / _compute_statistical_significance: check np.isnan(control).any() or np.isnan(treatment).any() before the permutation loop, and raise a ValueError naming the metric, or return (nan, False) with a warning. That way a NaN can never come out as p_value=0.0, significant=True. Apply the same check to the student path so an undefined result is reported as undefined.
ranx/metrics/bpref.py:13-30(_bpref):When a query's qrels contain no judged non-relevant documents (
n_non_rels == 0),min(n_rels, n_non_rels)is0. Every retrieved relevant document then contributes0/0, and the query's score isnan. Positive-only qrels are a common shape: MS MARCO-style qrels files, for example, list only relevant pairs. Under@njitno warning reaches the caller.This was reported in #66 and closed as expected behaviour ("with no known irrelevant documents you are dividing by zero … N is zero"). The maintainer also said ranx's bpref "should be the same as TREC Eval". trec_eval does not divide by zero here.
m_bpref.cspecial-cases it explicitly:With
num_nonrel == 0,nonrel_so_faris always0. Each retrieved relevant document adds1.0, and trec_eval returns (relevant retrieved) / R. On identical input, ranx returnsnan. This matches the #66 reporter's observation that trec_eval and pytrec_eval produced values for the same data.The NaN then hits a separate problem in
ranx/statistical_tests/fisher_randomization_test.py:29-46:If any per-query score is
nan, bothcontrol_treatment_diffand everypermuted_diffarenan, andnan >= nanisFalse.counter_arraytherefore never increments, and the test returnsp_value = 0.0, significant = True: the strongest possible "the systems differ" result, whatever the data. The bug is independent of bpref. Any NaN in the score arrays triggers it.Measured
Each case re-runs the function bodies copied verbatim from HEAD
7363db0c(v0.3.21) under numba 0.67 / numpy 2.5.3 / scipy 1.18.1, with@njit/parallel=Truekept. Doc IDs are small floats, as in ranx's float64 qrels/run arrays._bprefontests/unit/ranx/metrics/bpref_test.pyq1 (qrels{d2:1,d5:1,d1:0,d4:0}, rund1,d2,d3,d4) - control0.250.25(matches the test's assertion)np.isclose(0.25, 0.999)- must-fail controlFalseFalse[[2,1],[5,1],[7,1]](relevant only), run retrieves 2, 5, 7 first (perfect ranking) plus one unjudged docnan1.0(trec_eval: 3 relevant retrieved / R=3)nan0.333…(trec_eval: 1/3)0.00.0(no division reached)_bpref_parallelover [q1 fixture, case-3 query][0.25, nan], no warning[0.25, 1.0]n_rels == 0) - control0.00.0(trec_eval also returns 0)np.mean([nan, 0.5, 0.8, 0.9])(the aggregation used byevaluate()andcompare())nanfisher_randomization_test, control=[0.50,0.55,0.60,0.45,0.52,0.58,0.49,nan], treatment=[0.51,0.54,0.61,0.44,0.53,0.57,0.50,0.52], 1000 permutations,max_p=0.01p_value=0.0,significant=Truenanreplaced by0.51p_value=0.705,significant=False[0.0]*10vs[1.0]*10, 2000 permutationsp_value=0.0005,significant=Truepaired_student_t_testpath (ttest_relthenp_value <= max_p, line 15) on the case-9 arraysp_value=nan,significant=FalseConsequence
bprefnan.evaluate()then reports the run-levelbprefasnan(ranx/meta/evaluate.py:151,:162,:167, plainnp.mean), and so doescompare()(ranx/meta/compare.py:93,float(np.mean(...))), with nothing pointing at the query or the cause. trec_eval returns a defined score for the same input.compare(..., stat_test="fisher")passes the raw per-query arrays (compare.py:96) tofisher_randomization_test. For every pair of runs whose scores for that metric include anan, it reportsp_value: 0.0, significant: True, which makes two indistinguishable systems (case 9 vs. control case 10) look maximally significantly different. The same happens for any NaN-producing metric, not only bpref. (compare()defaults tostat_test="student", so this path needsstat_test="fisher"; the internalcompute_statistical_significancedefaults to"fisher".)studentpath the NaN becomessignificant: False. The return value cannot distinguish that from a genuine non-significant result.Suggested fix
_bpref: match trec_eval. A relevant document with no judged non-relevant document ranked above it contributes1.0, so the0/0never happens. For example, compute the per-relevant term only wherecumsum(hit_list == 0) > 0and use1.0otherwise, or return(relevant retrieved) / n_relswhenn_non_rels == 0. Add an explicitif n_rels == 0: return 0.0in the same style as the sibling guards (hits.py:13-15,recall.py:14-16,f1.py:14-16,average_precision.py:13-15,r_precision.py:16-17). Standalone@njit_bprefcurrently raisesZeroDivisionErrorin that case and relies on the parallel wrapper to yield0.0. If the current NaN behaviour is kept on purpose, a clear error or warning naming the query (as suggested in I'm getting NaN for the BPref measurement #66) would at least make it visible.fisher_randomization_test/_compute_statistical_significance: checknp.isnan(control).any() or np.isnan(treatment).any()before the permutation loop, and raise aValueErrornaming the metric, or return(nan, False)with a warning. That way a NaN can never come out asp_value=0.0, significant=True. Apply the same check to thestudentpath so an undefined result is reported as undefined.Happy to open the PR.