1 paper · 1 filter
David Méndez, Roberto Confalonieri, Natalia DÃaz RodrÃguez
Vision-Language Models (VLMs) excel at tasks like zero-shot classification and cross-modal retrieval by mapping images and text to a shared space, but this requires expensive end-t…