Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
M4-RAG: A Massive-Scale Multilingual Multi-Cultural Multimodal RAG
David Anugraha, Patrick Amadeus Irawan, Anshul Singh +2
Vision-language models (VLMs) have achieved strong performance in visual question answering (VQA), yet they remain constrained by static training data. Retrieval-Augmented Generati…
cs.CL2025
Lost in Translation and Noise: A Deep Dive into the Failure Modes of VLMs on Real-World Tables
Anshul Singh, Rohan Chaudhary, Gagneet Singh +1
The impressive performance of VLMs is largely measured on benchmarks that fail to capture the complexities of real-world scenarios. Existing datasets for tabular QA, such as WikiTa…