Hi OpenSearch-VL authors,
Thanks for releasing this excellent work.
I am trying to reproduce the evaluation results of OpenSearch-VL and would like to clarify the exact evaluation samples used in the experiments.Could you please provide the sample lists or split files used for the following benchmarks?
- FVQA evaluation:
- Which exact FVQA test samples were used for the reported results?
- Is the evaluation set the same as the FVQA test split from MMSearch-R1, or another split?
- LiveVQA evaluation:
- Which exact LiveVQA samples were used for evaluation?
- Since there are different versions/splits of LiveVQA, could you please specify the benchmark version and provide the corresponding sample IDs if possible?
Having the exact evaluation sample IDs would help ensure a faithful reproduction and facilitate comparison with future work.
Thank you very much for your help.
Hi OpenSearch-VL authors,
Thanks for releasing this excellent work.
I am trying to reproduce the evaluation results of OpenSearch-VL and would like to clarify the exact evaluation samples used in the experiments.Could you please provide the sample lists or split files used for the following benchmarks?
Having the exact evaluation sample IDs would help ensure a faithful reproduction and facilitate comparison with future work.
Thank you very much for your help.