Publications (9)
CountQA: How Well Do MLLMs Count in the Wild?
Jayant Sravan Tamarapalli, Rynaa Grover, Nilay Pande +1
Multimodal Large Language Models (MLLMs) demonstrate remarkable fluency in understanding visual scenes, yet they exhibit a critical lack in a fundamental cognitive skill: object co…
Database of Indian Social Media Influencers on Twitter
Arshia Arya, Soham De, Dibyendu Mishra +12
Databases of highly networked individuals have been indispensable in studying narratives and influence on social media. To support studies on Twitter in India, we present a systema…
Humanity's Last Exam
Long Phan, Alice Gatti, Ziwen Han +1144
Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achi…
Rihanna versus Bollywood: Twitter Influencers and the Indian Farmers' Protest
Dibyendu Mishra, Syeda Zainab Akbar, Arshia Arya +3
A tweet from popular entertainer and businesswoman, Rihanna, bringing attention to farmers' protests around Delhi set off heightened activity on Indian social media. An immediate c…
HueManity: Probing Fine-Grained Visual Perception in MLLMs
Rynaa Grover, Jayant Sravan Tamarapalli, Sahiti Yerramilli +1
Recent Multimodal Large Language Models (MLLMs) demonstrate strong high-level visual reasoning on tasks such as visual question answering and image captioning. Yet existing benchma…
GeoChain: Multimodal Chain-of-Thought for Geographic Reasoning
Sahiti Yerramilli, Nilay Pande, Rynaa Grover +1
This paper introduces GeoChain, a large-scale benchmark for evaluating step-by-step geographic reasoning in multimodal large language models (MLLMs). Leveraging 1.46 million Mapill…
A Community-Centric Perspective for Characterizing and Detecting Anti-Asian Violence-Provoking Speech
Gaurav Verma, Rynaa Grover, Jiawei Zhou +4
Violence-provoking speech -- speech that implicitly or explicitly promotes violence against the members of the targeted community, contributed to a massive surge in anti-Asian crim…
MaRVL-QA: A Benchmark for Mathematical Reasoning over Visual Landscapes
Nilay Pande, Sahiti Yerramilli, Jayant Sravan Tamarapalli +1
A key frontier for Multimodal Large Language Models (MLLMs) is the ability to perform deep mathematical and spatial reasoning directly from images, moving beyond their established…
Insights Into Incitement: A Computational Perspective on Dangerous Speech on Twitter in India
Saloni Dash, Rynaa Grover, Gazal Shekhawat +3
Dangerous speech on social media platforms can be framed as blatantly inflammatory, or be couched in innuendo. It is also centrally tied to who engages it - it can be driven by ope…