7 papers
Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models
Jiaang Li, Chengzu Li, Zhaochong An +4
Multimodal Large Language Models (MLLMs) achieve strong performance by integrating visual inputs with the rich priors of pretrained language models. However, they often fail on vis…
Thinking in Frames: How Visual Context and Test-Time Scaling Empower Video Reasoning
Chengzu Li, Zanyi Wang, Jiaang Li +9
Vision-Language Models have excelled at textual reasoning, but they often struggle with fine-grained spatial understanding and continuous action planning, failing to simulate the d…
What if Othello-Playing Language Models Could See?
Xinyi Chen, Yifei Yuan, Jiaang Li +3
Language models are often said to face a symbol grounding problem. While some have argued the problem can be solved without resort to other modalities, many have speculated that gr…
Evaluation of Cultural Competence of Vision-Language Models
Srishti Yadav, Lauren Tilton, Maria Antoniak +10
Modern vision-language models (VLMs) often fail at cultural competency evaluations and benchmarks. Given the diversity of applications built upon VLMs, there is renewed interest in…
RAVENEA: A Benchmark for Multimodal Retrieval-Augmented Visual Culture Understanding
Jiaang Li, Yifei Yuan, Wenyan Li +8
As vision-language models (VLMs) become increasingly integrated into daily life, the need for accurate visual culture understanding is becoming critical. Yet, these models frequent…
Multi-Modal Framing Analysis of News
Arnav Arora, Srishti Yadav, Maria Antoniak +2
Automated frame analysis of political communication is a popular task in computational social science that is used to study how authors select aspects of a topic to frame its recep…