Position Bias in Mamba and Hybrid Language Models | Evidence Position Bias Across Sequence Mixers in Long-Context Question Answering | NeurIPS'2026
-
Updated
Aug 22, 2026 - Python
Position Bias in Mamba and Hybrid Language Models | Evidence Position Bias Across Sequence Mixers in Long-Context Question Answering | NeurIPS'2026
ACL 2026 | CapCal: content-agnostic probability calibration for de-biasing listwise rerankers.
A reliability lab for LLM judges: measure position bias, verbosity bias, and calibration against human labels — then recalibrate. Pure stdlib; the demo needs no API keys.
position bias in LLM judges is worse than people think: ask a small instruct model to choose 1 of 2 items. 81% of the answer will be based on which slot the item was in, not which it was. swap the order and the answer flips 78% of the time. an inconclusive cognitive dissonance experiment
Position-bias-aware ranking: estimating and correcting position bias to optimize Earnings Per Visitor (EPV) · simulation study · IPW & propensity modeling · Python
Monthly bias audits of LLM judges
Ranking evaluation with error bars: NDCG, MRR and MAP with confidence intervals, plus position-bias correction for click logs. No dependencies.
Measure position/verbosity/assertiveness bias in an LLM-as-judge (Claude) by judging pairs in both orders. Finding: no position bias, but 75% verbosity bias and 100% assertiveness bias — judge scores gameable by length + tone.
Two kinds of saturation: why LLM-judge order bias is hard to measure — essay, Lean proofs (0 sorry), and a reproducible dispersion measurement.
Counterfactual learning-to-rank for marketplace search logs in PySpark: position-bias estimation, IPS-weighted training, NDCG evaluation against known ground truth
Add a description, image, and links to the position-bias topic page so that developers can more easily learn about it.
To associate your repository with the position-bias topic, visit your repo's landing page and select "manage topics."