activity
20242026
most citedEgocentric Bias in Vision-Language Models

1 citations · 1 across the 8 of their papers we have counts for

collaborators
Showing 2026Show all

6 papers · 1 filter

cs.CV2026

Self-Supervised Visual On-Policy Distillation

Yijiang Li, Yijun Liang, Yunjie Tian +6

Visual on-policy distillation relies heavily on an informative teacher-student asymmetry, through either a larger, stronger teacher or privileged supervision, such as reference ans…

cs.LG2026

On-Policy Self-Distillation without Any Supervision

Yijiang Li, Bingyang Wang, Yijun Liang +3

On-policy (Self-)Distillation (OPD / OPSD) has shown strong potential for post-training large language models (LLMs). However, existing methods still rely heavily on external super…

cs.AI2026

Contextualized Evaluation of Vision Language Models through Dynamic, Multi-turn Interactions

Yijiang Li, Huiqi Zou, Bingyang Wang +1

Multi-modal Large Language Models (MLLMs) have made substantial advances on benchmarks, yet their real-world effectiveness remains uncertain. This gap stems from the fundamental mi…

cs.AI2026

SPACENUM: Revisiting Spatial Numerical Understanding in VLMs

Jianshu Zhang, Yijiang Li, Huifeixin Chen +4

Vision-Language Models (VLMs) are increasingly deployed in embodied environments, where they need produce numerical outputs such as action magnitudes and spatial coordinates. Altho…

cs.AI2026

Vision Language Models Cannot Reason About Physical Transformation

Dezhi Luo, Yijiang Li, Maijunxian Wang +7

Understanding physical transformations is fundamental for reasoning in dynamic environments. While Vision Language Models (VLMs) show promise in embodied applications, whether they…

cs.CV2026★ 1 cited

Egocentric Bias in Vision-Language Models

Maijunxian Wang, Yijiang Li, Bingyang Wang +6

Visual perspective taking--inferring how the world appears from another's viewpoint--is foundational to social cognition. We introduce FlipSet, a diagnostic benchmark for Level-2 v…