1 paper
Jiajun Qin, Yuan Pu, Zhuolun He +3
Current vision-language models have been explored for multi-modal embedding tasks like information retrieval. However, they face significant challenges in real-world queries and ta…