Paola Cascante-Bonilla, Khaled Shehada, James Seale Smith +8
Large-scale pre-trained Vision & Language (VL) models have shown remarkable performance in many applications, enabling replacing a fixed set of supported classes with zero-shot ope…