2 citations · 2 across the 8 of their papers we have counts for
1 paper · 1 filter
Gesina Schwalbe, Mert Keser, Moritz Bayerkuhnlein +8
Vision-language model (VLM) encoders such as CLIP enable strong retrieval and zero-shot classification in a shared image-text embedding space, yet the semantic organization of this…