Official repository for “MTurn-Seg: A Large-Scale Bilingual Medical Benchmark for Multi-Turn Reasoning Segmentation (BIBM 2025)”. You can check the current version here: https://cowboyh.github.io/MTurn-Seg/. The official link to the paper: https://ieeexplore.ieee.org/abstract/document/11356076
Haitao Nie*, Yimeng Zheng*, Ying Ye*, Bin Xie†
Artificial Intelligence and Robotics Laboratory (AIRLab),
Central South University, Changsha, China
* Equal contribution.
† Corresponding author.
-
New Task — Multi-Turn Reasoning Segmentation (MTRS): At each turn, the model consumes the current instruction + interaction history (prior prompts and masks) to produce the next segmentation.
-
Three Reasoning Facets: (i) Clinical/Anatomical (e.g., “segment the solid organ in the right upper abdomen involved in glucose metabolism”), (ii) Spatial (e.g., “segment the elliptical structure adjacent to the right side of the abdominal aorta”), (iii) History-based References (e.g., “segment the necrotic region surrounding the previously segmented tumor”).
-
Bilingual Benchmark (ZH/EN): First dataset supporting multi-turn medical dialogues in Chinese and English.
-
Scale & Coverage: 28,904 images, 113,963 masks, 232,188 QA pairs across CT & MRI; covers major organs and anatomical systems.
-
What It Measures: Cross-turn memory, history-conditioned mask refinement, and language-to-image alignment over multiple rounds.
-
SOTA Evaluation: Benchmarked MedCLIP-SAM, LISA, and LISA++ under multi-turn settings.
-
Key Findings:
- Current models are well below clinical usability on this benchmark.
- Performance degrades as dialogue turns increase.
- General-purpose models outperform medical-specific models, indicating a need to infuse stronger domain knowledge.
-
Intended Impact: Establishes the first large-scale yardstick for MTRS, enabling fair, reproducible comparison and catalyzing progress on multi-turn reasoning in medical imaging.
MTurn-Seg Dataset License: CC BY-NC-SA 4.0
You are free to share and adapt the dataset under the following terms:
- Attribution (BY)
- NonCommercial (NC)
- ShareAlike (SA)
Full text: https://creativecommons.org/licenses/by-nc-sa/4.0/
If you find this work useful or use this dataset in your research, please cite our paper.
BibTeX
@INPROCEEDINGS{11356076,
author={Nie, Haitao and Zheng, Yimeng and Ye, Ying and Xie, Bin},
booktitle={2025 IEEE International Conference on Bioinformatics and Biomedicine (BIBM)},
title={MTurn-Seg: A Large-Scale Bilingual Medical Benchmark for Multi-Turn Reasoning Segmentation},
year={2025},
volume={},
number={},
pages={3946-3950},
keywords={Image segmentation;Pathology;Magnetic resonance imaging;Computed tomography;Medical services;Benchmark testing;Cognition;Usability;Standards;Biomedical imaging;Multi-Turn Medical Reasoning Segmentation;Bilingual;Benchmark},
doi={10.1109/BIBM66473.2025.11356076}}