From the 2 of 17 linked papers with an AI index.
1 citations · 1 across the 2 of their papers we have counts for
17 papers
Contextualized Evaluation of Vision Language Models through Dynamic, Multi-turn Interactions
Yijiang Li, Huiqi Zou, Bingyang Wang +1
The paper presents CEDI, a framework that evaluates vision‑language models through multi‑turn, interactive dialogues between the model, an automated examiner, and a grader, uncover…
Egocentric Bias in Vision-Language Models
Maijunxian Wang, Yijiang Li, Bingyang Wang +6
The paper introduces FlipSet, a benchmark that tests vision‑language models on Level‑2 visual perspective taking by requiring them to mentally rotate 2D character strings, and find…
ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model
Haichao Zhang, Yijiang Li, Shwai He +5
Recent progress in latent world models (e.g., V-JEPA2) has shown promising capability in forecasting future world states from video observations. Nevertheless, dense prediction fro…
Vision Language Models Cannot Reason About Physical Transformation
Dezhi Luo, Yijiang Li, Maijunxian Wang +7
Understanding physical transformations is fundamental for reasoning in dynamic environments. While Vision Language Models (VLMs) show promise in embodied applications, whether they…
Vision-Language Models Mistake Head Orientation for Gaze Direction: Nonverbal Conversation Cues
Zory Zhang, Pinyuan Feng, Bingyang Wang +7
Where someone looks is a nonverbal communication cue that children and adults readily use. How well can Vision-Language Models (VLMs) infer gaze targets? To construct evaluation st…
SPACENUM: Revisiting Spatial Numerical Understanding in VLMs
Jianshu Zhang, Yijiang Li, Huifeixin Chen +4
Vision-Language Models (VLMs) are increasingly deployed in embodied environments, where they need produce numerical outputs such as action magnitudes and spatial coordinates. Altho…