activity
20242026
collaborators

7 papers

cs.CV2026

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models

Jiaang Li, Chengzu Li, Zhaochong An +4

Multimodal Large Language Models (MLLMs) achieve strong performance by integrating visual inputs with the rich priors of pretrained language models. However, they often fail on vis…

cs.LG2026

Thinking in Frames: How Visual Context and Test-Time Scaling Empower Video Reasoning

Chengzu Li, Zanyi Wang, Jiaang Li +9

Vision-Language Models have excelled at textual reasoning, but they often struggle with fine-grained spatial understanding and continuous action planning, failing to simulate the d…

cs.AI2025

What if Othello-Playing Language Models Could See?

Xinyi Chen, Yifei Yuan, Jiaang Li +3

Language models are often said to face a symbol grounding problem. While some have argued the problem can be solved without resort to other modalities, many have speculated that gr…

cs.CV2025

Evaluation of Cultural Competence of Vision-Language Models

Srishti Yadav, Lauren Tilton, Maria Antoniak +10

Modern vision-language models (VLMs) often fail at cultural competency evaluations and benchmarks. Given the diversity of applications built upon VLMs, there is renewed interest in…

cs.CV2025

RAVENEA: A Benchmark for Multimodal Retrieval-Augmented Visual Culture Understanding

Jiaang Li, Yifei Yuan, Wenyan Li +8

As vision-language models (VLMs) become increasingly integrated into daily life, the need for accurate visual culture understanding is becoming critical. Yet, these models frequent…

cs.CL2025

Multi-Modal Framing Analysis of News

Arnav Arora, Srishti Yadav, Maria Antoniak +2

Automated frame analysis of political communication is a popular task in computational social science that is used to study how authors select aspects of a topic to frame its recep…