2 papers
cs.CV2026
How can embedding models bind concepts?
Arnas Uselis, Darina Koishigarina, Seong Joon Oh
Humans easily determine which color belongs to which shape in multi-object scenes, an ability known as concept binding. Vision-language embedding models such as CLIP struggle with…
cs.CV2026
CLIP Behaves like a Bag-of-Words Model Cross-modally but not Uni-modally
Darina Koishigarina, Arnas Uselis, Seong Joon Oh
CLIP (Contrastive Language-Image Pretraining) has become a popular choice for various downstream tasks. However, recent studies have questioned its ability to represent composition…