Official repository for the ICLR 2026 paper: "SportR: A Benchmark for Multimodal Large Language Model Reasoning in Sports".
- [2026.03] 🎉 SportR has been accepted to ICLR 2026! We are currently finalizing the dataset for public release.
SportR is the large-scale, multi-sport benchmark designed to evaluate and train Multimodal Large Language Models (MLLMs) on complex reasoning. It challenges models to master three core capabilities: visual perception, sport-rule knowledge, and visual grounding.
- Multi-Modal: 4,789 images (SportsImage) and 2,052 videos (SportsVideo) with over 20,000 question-answer pairs.
- Complex Reasoning: 6,841 high-quality, human-authored rationale annotations.
- Fine-grained Grounding: bounding box annotations for Images.
- Diverse Sports: Comprehensive coverage across multiple sporting disciplines and infraction types.
We are releasing the benchmark in stages:
- Phase 1: Project page launch.
- Phase 2: Dataset released — QA + CoT annotations in
annotations/; images & videos available on 🤗 Hugging Face.
If you find our work helpful for your research, please cite:
@misc{xia2025sportrbenchmarkmultimodallarge,
title={SportR: A Benchmark for Multimodal Large Language Model Reasoning in Sports},
author={Haotian Xia and Haonan Ge and Junbo Zou and Hyun Woo Choi and Xuebin Zhang and Danny Suradja and Botao Rui and Ethan Tran and Wendy Jin and Zhen Ye and Xiyang Lin and Christopher Lai and Shengjie Zhang and Junwen Miao and Shichao Chen and Rhys Tracy and Vicente Ordonez and Weining Shen and Hanjie Chen},
year={2025},
eprint={2511.06499},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2511.06499},
}