works on

From the 2 of 17 linked papers with an AI index.

most citedEgocentric Bias in Vision-Language Models

1 citations · 1 across the 2 of their papers we have counts for

collaborators

17 papers

cs.AI2026

Contextualized Evaluation of Vision Language Models through Dynamic, Multi-turn Interactions

Yijiang Li, Huiqi Zou, Bingyang Wang +1

The paper presents CEDI, a framework that evaluates vision‑language models through multi‑turn, interactive dialogues between the model, an automated examiner, and a grader, uncover…

cs.CV20261 cited

Egocentric Bias in Vision-Language Models

Maijunxian Wang, Yijiang Li, Bingyang Wang +6

The paper introduces FlipSet, a benchmark that tests vision‑language models on Level‑2 visual perspective taking by requiring them to mentally rotate 2D character strings, and find…

cs.CV2026

ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model

Haichao Zhang, Yijiang Li, Shwai He +5

Recent progress in latent world models (e.g., V-JEPA2) has shown promising capability in forecasting future world states from video observations. Nevertheless, dense prediction fro…

cs.AI2026

Vision Language Models Cannot Reason About Physical Transformation

Dezhi Luo, Yijiang Li, Maijunxian Wang +7

Understanding physical transformations is fundamental for reasoning in dynamic environments. While Vision Language Models (VLMs) show promise in embodied applications, whether they…

cs.CV2026

Vision-Language Models Mistake Head Orientation for Gaze Direction: Nonverbal Conversation Cues

Zory Zhang, Pinyuan Feng, Bingyang Wang +7

Where someone looks is a nonverbal communication cue that children and adults readily use. How well can Vision-Language Models (VLMs) infer gaze targets? To construct evaluation st…

cs.AI2026

SPACENUM: Revisiting Spatial Numerical Understanding in VLMs

Jianshu Zhang, Yijiang Li, Huifeixin Chen +4

Vision-Language Models (VLMs) are increasingly deployed in embodied environments, where they need produce numerical outputs such as action magnitudes and spatial coordinates. Altho…