Monte Carlo bias probe for LLM-as-judge evals — measures position, verbosity, and formatting bias, and how much swap-averaging removes.
nlp machine-learning claude prompt-engineering anthropic llm-evaluation llm-as-judge bias-monte-carlo
-
Updated
Jun 18, 2026 - Python