1 paper
Nidham Tekaya, Manuela Waldner, Matthias Zeppelzauer
Large-scale vision-language models (VLMs) such as CLIP have gained popularity for their generalizable and expressive multimodal representations. By leveraging large-scale training…