Hi authors,
Thanks for the excellent work and I am very interested about this paper! I am trying to reproduce it, especially the training part. I am curious about the typical range of the cosine similarity loss, since in my experiment the alignment metric peaks above 0.9 during early iterations, then remains on a long plateau, which seems very strange. Could you share what range the alignment loss typically lies in and whether this behavior is expected?
Hi authors,
Thanks for the excellent work and I am very interested about this paper! I am trying to reproduce it, especially the training part. I am curious about the typical range of the cosine similarity loss, since in my experiment the alignment metric peaks above 0.9 during early iterations, then remains on a long plateau, which seems very strange. Could you share what range the alignment loss typically lies in and whether this behavior is expected?