1 paper · 1 filter
Wenhao Wang, Adam Dziedzic, Grace C. Kim +2
Multi-modal models, such as CLIP, have demonstrated strong performance in aligning visual and textual representations, excelling in tasks like image retrieval and zero-shot classif…