To adapt the corpus to our setting, for each of the 32K test images, we sample a (FOIL, true) pair, and compute the accuracy of each evaluation metric in their capacity to assign a higher score to the true candidate versus the FOIL.
Hi there, I want to reproduce Table 5(corresponding to section 4.4 Sensitivity of CLIP-S to hallucination), and found there is some randomness for sampling (FOIL) dataset. Do you happen to have any log of sampled image key to reproduce same result?
Hi there, I want to reproduce Table 5(corresponding to section 4.4 Sensitivity of CLIP-S to hallucination), and found there is some randomness for sampling (FOIL) dataset. Do you happen to have any log of sampled image key to reproduce same result?