1 paper
Yunhao Ge, Jie Ren, Andrew Gallagher +6
Multi-modal image-text models such as CLIP and LiT have demonstrated impressive performance on image classification benchmarks and their zero-shot generalization ability is particu…