From the 1 of 8 linked papers with an AI index.
8 papers
Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models
Jiaang Li, Chengzu Li, Zhaochong An +5
The paper investigates why multimodal large language models often ignore visual evidence, using image reconstruction and a new benchmark (WhatIfVis) to measure how well models bala…
RAVENEA: A Benchmark for Multimodal Retrieval-Augmented Visual Culture Understanding
Jiaang Li, Yifei Yuan, Wenyan Li +8
As vision-language models (VLMs) become increasingly integrated into daily life, the need for accurate visual culture understanding is becoming critical. Yet, these models frequent…
Thinking in Frames: How Visual Context and Test-Time Scaling Empower Video Reasoning
Chengzu Li, Zanyi Wang, Jiaang Li +9
Vision-Language Models have excelled at textual reasoning, but they often struggle with fine-grained spatial understanding and continuous action planning, failing to simulate the d…
What if Othello-Playing Language Models Could See?
Xinyi Chen, Yifei Yuan, Jiaang Li +3
Language models are often said to face a symbol grounding problem. While some have argued the problem can be solved without resort to other modalities, many have speculated that gr…
Revealing Fine-Grained Values and Opinions in Large Language Models
Dustin Wright, Arnav Arora, Nadav Borenstein +3
Uncovering latent values and opinions embedded in large language models (LLMs) can help identify biases and mitigate potential harm. Recently, this has been approached by prompting…
Evaluation of Cultural Competence of Vision-Language Models
Srishti Yadav, Lauren Tilton, Maria Antoniak +10
Modern vision-language models (VLMs) often fail at cultural competency evaluations and benchmarks. Given the diversity of applications built upon VLMs, there is renewed interest in…