11 papers
Compositional Generalization Requires Linear, Orthogonal Representations in Vision Embedding Models
Arnas Uselis, Andrea Dittadi, Seong Joon Oh
Compositional generalization, the ability to recognize familiar parts in novel contexts, is a defining property of intelligent systems. Although modern models are trained on massiv…
When Do Diffusion Models learn to Generate Multiple Objects?
Yujin Jeong, Arnas Uselis, Iro Laina +2
Text-to-image diffusion models achieve impressive visual fidelity, yet they remain unreliable in multi-object generation. Despite extensive empirical evidence of these failures, th…
How can embedding models bind concepts?
Arnas Uselis, Darina Koishigarina, Seong Joon Oh
Humans easily determine which color belongs to which shape in multi-object scenes, an ability known as concept binding. Vision-language embedding models such as CLIP struggle with…
Sparse Autoencoders enable Robust and Interpretable Fine-tuning of CLIP models
Fabian Morelli, Arnas Uselis, Ankit Sonthalia +1
Large-scale pre-trained vision-language models like CLIP demonstrate remarkable zero-shot performance across diverse tasks. However, fine-tuning these models to improve downstream…
MEME: Multi-entity & Evolving Memory Evaluation
Seokwon Jung, Alexander Rubinstein, Arnas Uselis +2
LLM-based agents increasingly operate in persistent environments where they must store, update, and reason over information across many sessions. While prior benchmarks evaluate on…
CLIP Behaves like a Bag-of-Words Model Cross-modally but not Uni-modally
Darina Koishigarina, Arnas Uselis, Seong Joon Oh
CLIP (Contrastive Language-Image Pretraining) has become a popular choice for various downstream tasks. However, recent studies have questioned its ability to represent composition…