1 paper · 1 filter
Anton Baumann, Rui Li, Marcus Klasson +5
Vision-language models (VLMs), such as CLIP and SigLIP, have found remarkable success in classification, retrieval, and generative tasks. For this, VLMs deterministically map image…