1 paper
Tingyu Song, Mingxin Li, Yanzhao Zhang +5
MLLM-based embedding models remain limited in compositional retrieval, often failing to distinguish scenes containing the same concepts but different attribute-object bindings. Yet…