Showing cs.AIShow all
3 papers · 1 filter
cs.AI2025
MaRVL-QA: A Benchmark for Mathematical Reasoning over Visual Landscapes
Nilay Pande, Sahiti Yerramilli, Jayant Sravan Tamarapalli +1
A key frontier for Multimodal Large Language Models (MLLMs) is the ability to perform deep mathematical and spatial reasoning directly from images, moving beyond their established…
cs.AI2025
CountQA: How Well Do MLLMs Count in the Wild?
Jayant Sravan Tamarapalli, Rynaa Grover, Nilay Pande +1
Multimodal Large Language Models (MLLMs) demonstrate remarkable fluency in understanding visual scenes, yet they exhibit a critical lack in a fundamental cognitive skill: object co…
cs.AI2025
GeoChain: Multimodal Chain-of-Thought for Geographic Reasoning
Sahiti Yerramilli, Nilay Pande, Rynaa Grover +1
This paper introduces GeoChain, a large-scale benchmark for evaluating step-by-step geographic reasoning in multimodal large language models (MLLMs). Leveraging 1.46 million Mapill…