I reproduced the training process of Prime Verifier using your code and obtained curves similar to those in the paper. However, when I used the trained model to generate answers on the dataset and evaluated them, the measured accuracy differed noticeably from the accuracies shown in those curves—especially on the Olympiad_bench dataset. What could be causing this discrepancy?

I reproduced the training process of Prime Verifier using your code and obtained curves similar to those in the paper. However, when I used the trained model to generate answers on the dataset and evaluated them, the measured accuracy differed noticeably from the accuracies shown in those curves—especially on the Olympiad_bench dataset. What could be causing this discrepancy?