1 paper · 1 filter
Abrar Fahim, Alex Murphy, Alona Fyshe
Multi-modal contrastive models such as CLIP achieve state-of-the-art performance in zero-shot classification by embedding input images and texts on a joint representational space.…