From the 2 of 8 linked papers with an AI index.
1 citations · 1 across the 3 of their papers we have counts for
8 papers
On-Policy Self-Distillation without Any Supervision
Yijiang Li, Bingyang Wang, Yijun Liang +3
On-policy (Self-)Distillation (OPD / OPSD) has shown strong potential for post-training large language models (LLMs). However, existing methods still rely heavily on external super…
Contextualized Evaluation of Vision Language Models through Dynamic, Multi-turn Interactions
Yijiang Li, Huiqi Zou, Bingyang Wang +1
The paper presents CEDI, a framework that evaluates vision‑language models through multi‑turn, interactive dialogues between the model, an automated examiner, and a grader, uncover…
Egocentric Bias in Vision-Language Models
Maijunxian Wang, Yijiang Li, Bingyang Wang +6
The paper introduces FlipSet, a benchmark that tests vision‑language models on Level‑2 visual perspective taking by requiring them to mentally rotate 2D character strings, and find…
Vision Language Models Cannot Reason About Physical Transformation
Dezhi Luo, Yijiang Li, Maijunxian Wang +7
Understanding physical transformations is fundamental for reasoning in dynamic environments. While Vision Language Models (VLMs) show promise in embodied applications, whether they…
Vision-Language Models Mistake Head Orientation for Gaze Direction: Nonverbal Conversation Cues
Zory Zhang, Pinyuan Feng, Bingyang Wang +7
Where someone looks is a nonverbal communication cue that children and adults readily use. How well can Vision-Language Models (VLMs) infer gaze targets? To construct evaluation st…
SPACENUM: Revisiting Spatial Numerical Understanding in VLMs
Jianshu Zhang, Yijiang Li, Huifeixin Chen +4
Vision-Language Models (VLMs) are increasingly deployed in embodied environments, where they need produce numerical outputs such as action magnitudes and spatial coordinates. Altho…