1 paper
Anh-Quan Cao, Maximilian Jaritz, Matthieu Guillaumin +2
Large-scale vision-language pre-trained (VLP) models (e.g., CLIP) are renowned for their versatility, as they can be applied to diverse applications in a zero-shot setup. However,…