1 paper
Mayug Maniparambil, Chris Vorster, Derek Molloy +3
Contrastive pretrained large Vision-Language Models (VLMs) like CLIP have revolutionized visual representation learning by providing good performance on downstream datasets. VLMs a…