5 citations · 6 across the 3 of their papers we have counts for
1 paper · 1 filter
Haonan Chen, Hong Liu, Yuping Luo +4
Multimodal embedding models, built upon causal Vision Language Models (VLMs), have shown promise in various tasks. However, current approaches face three key limitations: the use o…