2 papers
cs.CV2024
Modeling Caption Diversity in Contrastive Vision-Language Pretraining
Samuel Lavoie, Polina Kirichenko, Mark Ibrahim +4
There are a thousand ways to caption an image. Contrastive Language Pretraining (CLIP) on the other hand, works by mapping an image and its caption to a single vector -- limiting h…
cs.CV2023
Understanding the Detrimental Class-level Effects of Data Augmentation
Polina Kirichenko, Mark Ibrahim, Randall Balestriero +4
Data augmentation (DA) encodes invariance and provides implicit regularization critical to a model's performance in image classification tasks. However, while DA improves average a…