1 paper
Benjamin Devillers, Bhavin Choksi, Romain Bielawski +1
Vision models trained on multimodal datasets can benefit from the wide availability of large image-caption datasets. A recent model (CLIP) was found to generalize well in zero-shot…