7 papers
Argus-Retriever: Vision-LLM Late-Interaction Retrieval with Region-Aware Query-Conditioned MoE for Visual Document Retrieval
Abdelrahman Abdallah, Mahmoud Abdalla, Mohammed Ali +1
Late-interaction vision-language retrievers represent each document page as many visual token embeddings and score queries with MaxSim. In systems such as ColPali, ColQwen, ColNomi…
HIVE: Query, Hypothesize, Verify An LLM Framework for Multimodal Reasoning-Intensive Retrieval
Mahmoud Abdalla, Mahmoud SalahEldin Kasem, Mohamed Mahmoud +3
Multimodal retrieval models fail on reasoning-intensive queries where images (diagrams, charts, screenshots) must be deeply integrated with text to identify relevant documents -- t…
BRIDGE: Multimodal-to-Text Retrieval via Reinforcement-Learned Query Alignment
Mohamed Darwish Mounis, Mohamed Mahmoud, Shaimaa Sedek +4
Multimodal retrieval systems struggle to resolve image-text queries against text-only corpora: the best vision-language encoder achieves only 27.6 nDCG@10 on MM-BRIGHT, underperfor…
MARVEL: Multimodal Adaptive Reasoning-intensiVe Expand-rerank and retrievaL
Mahmoud SalahEldin Kasem, Mohamed Mahmoud, Mostafa Farouk Senussi +3
Multimodal retrieval over text corpora remains a fundamental challenge: the best vision-language encoder achieves only 27.6 nDCG@10 on MM-BRIGHT, a reasoning-intensive multimodal r…
MM-BRIGHT: A Multi-Task Multimodal Benchmark for Reasoning-Intensive Retrieval
Abdelrahman Abdallah, Mohamed Darwish Mounis, Mahmoud Abdalla +6
Existing retrieval benchmarks primarily consist of text-based queries where keyword or semantic matching is usually sufficient. Many real-world queries contain multimodal elements,…
ReceiptSense: Beyond Traditional OCR -- A Dataset for Receipt Understanding
Abdelrahman Abdallah, Mohamed Mounis, Mahmoud Abdalla +6
Multilingual OCR and information extraction from receipts remains challenging, particularly for complex scripts like Arabic. We introduce \dataset, a comprehensive dataset designed…