Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
CLIP Behaves like a Bag-of-Words Model Cross-modally but not Uni-modally
Darina Koishigarina, Arnas Uselis, Seong Joon Oh
CLIP (Contrastive Language-Image Pretraining) has become a popular choice for various downstream tasks. However, recent studies have questioned its ability to represent composition…
cs.CV2025
Diffusion Classifiers Understand Compositionality, but Conditions Apply
Yujin Jeong, Arnas Uselis, Seong Joon Oh +1
Understanding visual scenes is fundamental to human intelligence. While discriminative models have significantly advanced computer vision, they often struggle with compositional un…
cs.CV2025
On the rankability of visual embeddings
Ankit Sonthalia, Arnas Uselis, Seong Joon Oh
We study whether visual embedding models capture continuous, ordinal attributes along linear directions, which we term _rank axes_. We define a model as _rankable_ for an attribute…