Skip to content

Fisher's Randomization Test yields nondeterministic results #70

Description

@marianrh

Describe the bug
When running Fisher's Randomization Test multiple times using the same data and same parameters, the significance results vary between runs. This happens even when the random seed is fixed and also when just a single thread is used.

To Reproduce
Run ranx.compare(qrels, runs, metrics=["precision@1", "recall@20"], stat_test="fisher", max_p=0.05, n_permutations=1000, make_comparable=True, threads=1, random_seed=0).
qrels has a few thousand entries, each element in runs about a thousand entries.

When running compare() multiple times, the significance assessments are slightly different between runs. On my data, this can be observed after three to five compare() runs.

Expected behavior
I expect multiple compare() runs on the same data and with the same parameters to always show the exact same significance results.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions