From the 1 of 10 linked papers with an AI index.
10 papers
Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models
Jiaang Li, Chengzu Li, Zhaochong An +4
The paper investigates why multimodal large language models often ignore visual evidence, using image reconstruction and a new benchmark (WhatIfVis) to measure how well models bala…
Video Understanding: From Geometry and Semantics to Unified Models
Zhaochong An, Zirui Li, Mingqiao Ye +9
Video understanding aims to enable models to perceive, reason about, and interact with the dynamic visual world. In contrast to image understanding, video understanding inherently…
RAVENEA: A Benchmark for Multimodal Retrieval-Augmented Visual Culture Understanding
Jiaang Li, Yifei Yuan, Wenyan Li +8
As vision-language models (VLMs) become increasingly integrated into daily life, the need for accurate visual culture understanding is becoming critical. Yet, these models frequent…
Thinking in Frames: How Visual Context and Test-Time Scaling Empower Video Reasoning
Chengzu Li, Zanyi Wang, Jiaang Li +9
Vision-Language Models have excelled at textual reasoning, but they often struggle with fine-grained spatial understanding and continuous action planning, failing to simulate the d…
EvalCards: A Framework for Standardized Evaluation Reporting
Ruchira Dhar, Danae Sanchez Villegas, Antonia Karamolegkou +11
Evaluation has long been a central concern in NLP, and transparent reporting practices are more critical than ever in today's landscape of rapidly released open-access models. Draw…
What if Othello-Playing Language Models Could See?
Xinyi Chen, Yifei Yuan, Jiaang Li +3
Language models are often said to face a symbol grounding problem. While some have argued the problem can be solved without resort to other modalities, many have speculated that gr…