1 paper
Gesina Schwalbe, Mert Keser, Moritz Bayerkuhnlein +9
Vision-language model (VLM) encoders such as CLIP enable strong retrieval and zero-shot classification in a shared image-text embedding space, yet the semantic organization of this…