A live, contamination-resistant benchmark for evaluating LLM agents on U.S. macroeconomic nowcasting.
-
Updated
Aug 29, 2026 - Python
A live, contamination-resistant benchmark for evaluating LLM agents on U.S. macroeconomic nowcasting.
WC2026-Agents: Claude, ChatGPT (GPT-5.5), Gemini and Grok forecast and bet on all 104 FIFA World Cup 2026 matches vs real betting odds. Open data + code. arXiv:2607.17765
Inverted generation: construct the label first, render the sentence second. Strictly-labelled, contamination-free synthetic data without a judge.
A deduction benchmark that verifies itself: every instance provably unique, provably solvable without guessing, machine-graded.
Predict how a policy games a reward. RL environments scored by an executable-exploit verifier.
Do models catch flawed ML results? Flawed reports with byte-identical matched controls.
ML training defects that never crash. An agentic benchmark with patch-and-rerun, unit-test verification.
To associate your repository with the contamination-free topic, visit your repo's landing page and select "manage topics."