1 paper
Mohamed Darwish Mounis, Mohamed Mahmoud, Shaimaa Sedek +4
Multimodal retrieval systems struggle to resolve image-text queries against text-only corpora: the best vision-language encoder achieves only 27.6 nDCG@10 on MM-BRIGHT, underperfor…