Accepted to the ICLR 2023 TrustML-(un)Limited workshop
Pre-trained networks, esp. those trained on ImageNet, are widely used in Computer Vision. However, watermarks in ImageNet images can lead to learning artifacts in pre-trained networks. In this paper, we assess the extent of this behavior in popular pre-trained models and identify affected classes. Our analysis shows that multiple ImageNet classes, not just the ``carton'' class, rely on spurious correlations with watermarks. We propose a simple approach to mitigate this issue in fine-tuned networks by ignoring the most watermark-sensitive encodings.
This GitHub repository includes multiple notebooks related to a research paper.
- Dataset Generation & Collection of Activations The first notebook, "Dataset Generation & Collection of Activations," includes code for generating probing datasets and collecting activations from popular ImageNet pretrained models.
- Analysis The second notebook, "Analysis," contains code for performing experiments to evaluate the differentiability of output logit representations and analyzing feature extractor representations.
- Ignoring sensitive embeddings The third notebook, "Ignoring Sensitive Embeddings," includes two sub-notebooks.
- Training The "Training" sub-notebook provides code for training models on the Cal-Tech 256 image classification problem, fine-tuned on DenseNet-161 features. The "Analysis" sub-notebook includes code for analyzing the fine-tuned networks.
- Analysis The "Analysis" sub-notebook includes code for analyzing the fine-tuned networks.
