In the paper, the method uses class labels or image caption generated by clipcap etc. to composite a text to form text embedding while the actual way in the code doesn't include only hard-coded text and no extra caption or imagenet label. I wonder whether the code is the actual training code used to produce your pretrained model. By the way, the performance is great. I guess maybe the extra training data from Danbooru2021 takes effect but I'm not sure. Thanks for answering.
In the paper, the method uses class labels or image caption generated by clipcap etc. to composite a text to form text embedding while the actual way in the code doesn't include only hard-coded text and no extra caption or imagenet label. I wonder whether the code is the actual training code used to produce your pretrained model. By the way, the performance is great. I guess maybe the extra training data from Danbooru2021 takes effect but I'm not sure. Thanks for answering.