For the grounding task, does MedMO output coordinates normalized to 0–1000 like Qwen3-VL, or does it use absolute image dimensions?
Also, why is the performance of MedMO-Next4B on MedSG-Bench very poor in my experiments, far below the results reported in the paper?
If possible, could you please share the prompt you used for this task?
this is normalized to 0–1000 result table:
| task |
samples |
mean IoU |
IoU@0.1 recall |
IoU@0.3 recall |
IoU@0.5 recall |
valid parse |
| multi_view |
1148 |
0.1118646637 |
0.2177700348 |
0.1611498258 |
0.1097560976 |
0.3928571429 |
| object_tracking |
1000 |
0.0899920427 |
0.2000000000 |
0.1270000000 |
0.0760000000 |
0.3360000000 |
| referring |
1482 |
0.2232926066 |
|
|
|
|
task samples mean IoU IoU@0.1 recall IoU@0.3 recall IoU@0.5 recall valid parse
multi_view 1148 0.1118646637 0.2177700348 0.1611498258 0.1097560976 0.3928571429
object_tracking 1000 0.0899920427 0.2000000000 0.1270000000 0.0760000000 0.3360000000
referring 1482 0.2232926066
For the grounding task, does MedMO output coordinates normalized to 0–1000 like Qwen3-VL, or does it use absolute image dimensions?
Also, why is the performance of MedMO-Next4B on MedSG-Bench very poor in my experiments, far below the results reported in the paper?
If possible, could you please share the prompt you used for this task?
this is normalized to 0–1000 result table:
task samples mean IoU IoU@0.1 recall IoU@0.3 recall IoU@0.5 recall valid parse
multi_view 1148 0.1118646637 0.2177700348 0.1611498258 0.1097560976 0.3928571429
object_tracking 1000 0.0899920427 0.2000000000 0.1270000000 0.0760000000 0.3360000000
referring 1482 0.2232926066