3 papers
cs.IR2026
HIVE: Query, Hypothesize, Verify An LLM Framework for Multimodal Reasoning-Intensive Retrieval
Mahmoud Abdalla, Mahmoud SalahEldin Kasem, Mohamed Mahmoud +3
Multimodal retrieval models fail on reasoning-intensive queries where images (diagrams, charts, screenshots) must be deeply integrated with text to identify relevant documents -- t…
cs.IR2026
BRIDGE: Multimodal-to-Text Retrieval via Reinforcement-Learned Query Alignment
Mohamed Darwish Mounis, Mohamed Mahmoud, Shaimaa Sedek +4
Multimodal retrieval systems struggle to resolve image-text queries against text-only corpora: the best vision-language encoder achieves only 27.6 nDCG@10 on MM-BRIGHT, underperfor…
cs.IR2026
MARVEL: Multimodal Adaptive Reasoning-intensiVe Expand-rerank and retrievaL
Mahmoud SalahEldin Kasem, Mohamed Mahmoud, Mostafa Farouk Senussi +3
Multimodal retrieval over text corpora remains a fundamental challenge: the best vision-language encoder achieves only 27.6 nDCG@10 on MM-BRIGHT, a reasoning-intensive multimodal r…