Skip to content

bpref returns NaN for qrels with no judged non-relevant docs (trec_eval returns a defined value), and fisher_randomization_test reports p_value=0.0 / significant=True whenever a per-query score is NaN #85

Description

@shaurya416

ranx/metrics/bpref.py:13-30 (_bpref):

def _bpref(qrels, run, rel_lvl):
    n_rels = sum(qrels[:, 1] >= rel_lvl)
    n_non_rels = sum(qrels[:, 1] < rel_lvl)

    hit_list = _get_hit_list(qrels, run, k=0, rel_lvl=rel_lvl)
    unjudged_list = _get_unjudged_list(qrels, run, k=0)

    # Remove unjudged from hit_list
    hit_list = hit_list[unjudged_list != 1]

    return (
        np.sum(
            1.0
            - np.minimum(np.cumsum(hit_list == 0)[hit_list == 1], n_rels)
            / min(n_rels, n_non_rels)
        )
        / n_rels
    )

When a query's qrels contain no judged non-relevant documents (n_non_rels == 0), min(n_rels, n_non_rels) is 0. Every retrieved relevant document then contributes 0/0, and the query's score is nan. Positive-only qrels are a common shape: MS MARCO-style qrels files, for example, list only relevant pairs. Under @njit no warning reaches the caller.

This was reported in #66 and closed as expected behaviour ("with no known irrelevant documents you are dividing by zero … N is zero"). The maintainer also said ranx's bpref "should be the same as TREC Eval". trec_eval does not divide by zero here. m_bpref.c special-cases it explicitly:

/* Special case nonrel_so_far == 0 to avoid division by 0 */
if (nonrel_so_far > 0) {
    bpref += 1.0 -
        (((double) MIN(nonrel_so_far, res_rels.num_rel)) /
         (double) MIN(num_nonrel, res_rels.num_rel));
} else
    bpref += 1.0;
...
if (res_rels.num_rel)
    bpref /= res_rels.num_rel;

With num_nonrel == 0, nonrel_so_far is always 0. Each retrieved relevant document adds 1.0, and trec_eval returns (relevant retrieved) / R. On identical input, ranx returns nan. This matches the #66 reporter's observation that trec_eval and pytrec_eval produced values for the same data.

The NaN then hits a separate problem in ranx/statistical_tests/fisher_randomization_test.py:29-46:

control_treatment_diff = abs(control_mean - treatment_mean)
...
for i in prange(n_permutations):
    permuted = permute(control_treatment_stack)
    permuted_diff = abs(permuted[:, 0].mean() - permuted[:, 1].mean())
    if permuted_diff >= control_treatment_diff:
        counter_array[i] = 1.0
p_value = counter_array.mean()
return p_value, p_value <= max_p

If any per-query score is nan, both control_treatment_diff and every permuted_diff are nan, and nan >= nan is False. counter_array therefore never increments, and the test returns p_value = 0.0, significant = True: the strongest possible "the systems differ" result, whatever the data. The bug is independent of bpref. Any NaN in the score arrays triggers it.

Measured

Each case re-runs the function bodies copied verbatim from HEAD 7363db0c (v0.3.21) under numba 0.67 / numpy 2.5.3 / scipy 1.18.1, with @njit / parallel=True kept. Doc IDs are small floats, as in ranx's float64 qrels/run arrays.

# Case Result Expected
1 _bpref on tests/unit/ranx/metrics/bpref_test.py q1 (qrels {d2:1,d5:1,d1:0,d4:0}, run d1,d2,d3,d4) - control 0.25 0.25 (matches the test's assertion)
2 np.isclose(0.25, 0.999) - must-fail control False False
3 qrels [[2,1],[5,1],[7,1]] (relevant only), run retrieves 2, 5, 7 first (perfect ranking) plus one unjudged doc nan 1.0 (trec_eval: 3 relevant retrieved / R=3)
4 Same qrels, run = one unjudged doc, then doc 2 nan 0.333… (trec_eval: 1/3)
5 Same qrels, run retrieves no relevant doc 0.0 0.0 (no division reached)
6 Public path _bpref_parallel over [q1 fixture, case-3 query] [0.25, nan], no warning [0.25, 1.0]
7 Public path, query with only non-relevant judgments (n_rels == 0) - control 0.0 0.0 (trec_eval also returns 0)
8 np.mean([nan, 0.5, 0.8, 0.9]) (the aggregation used by evaluate() and compare()) nan a per-query NaN should not silently poison the run mean
9 fisher_randomization_test, control=[0.50,0.55,0.60,0.45,0.52,0.58,0.49,nan], treatment=[0.51,0.54,0.61,0.44,0.53,0.57,0.50,0.52], 1000 permutations, max_p=0.01 p_value=0.0, significant=True an error or a NaN-flagged result, not maximal significance
10 Control: same arrays with the nan replaced by 0.51 p_value=0.705, significant=False not significant
11 Control: [0.0]*10 vs [1.0]*10, 2000 permutations p_value=0.0005, significant=True significant
12 paired_student_t_test path (ttest_rel then p_value <= max_p, line 15) on the case-9 arrays p_value=nan, significant=False distinguishable from "not significant"

Consequence

  • One query whose qrels list only relevant documents makes that query's bpref nan. evaluate() then reports the run-level bpref as nan (ranx/meta/evaluate.py:151, :162, :167, plain np.mean), and so does compare() (ranx/meta/compare.py:93, float(np.mean(...))), with nothing pointing at the query or the cause. trec_eval returns a defined score for the same input.
  • compare(..., stat_test="fisher") passes the raw per-query arrays (compare.py:96) to fisher_randomization_test. For every pair of runs whose scores for that metric include a nan, it reports p_value: 0.0, significant: True, which makes two indistinguishable systems (case 9 vs. control case 10) look maximally significantly different. The same happens for any NaN-producing metric, not only bpref. (compare() defaults to stat_test="student", so this path needs stat_test="fisher"; the internal compute_statistical_significance defaults to "fisher".)
  • On the default student path the NaN becomes significant: False. The return value cannot distinguish that from a genuine non-significant result.

Suggested fix

  1. _bpref: match trec_eval. A relevant document with no judged non-relevant document ranked above it contributes 1.0, so the 0/0 never happens. For example, compute the per-relevant term only where cumsum(hit_list == 0) > 0 and use 1.0 otherwise, or return (relevant retrieved) / n_rels when n_non_rels == 0. Add an explicit if n_rels == 0: return 0.0 in the same style as the sibling guards (hits.py:13-15, recall.py:14-16, f1.py:14-16, average_precision.py:13-15, r_precision.py:16-17). Standalone @njit _bpref currently raises ZeroDivisionError in that case and relies on the parallel wrapper to yield 0.0. If the current NaN behaviour is kept on purpose, a clear error or warning naming the query (as suggested in I'm getting NaN for the BPref measurement #66) would at least make it visible.
  2. fisher_randomization_test / _compute_statistical_significance: check np.isnan(control).any() or np.isnan(treatment).any() before the permutation loop, and raise a ValueError naming the metric, or return (nan, False) with a warning. That way a NaN can never come out as p_value=0.0, significant=True. Apply the same check to the student path so an undefined result is reported as undefined.

Happy to open the PR.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions