From the 1 of 7 linked papers with an AI index.
1 citations · 1 across the 1 of their papers we have counts for
7 papers
Egocentric Bias in Vision-Language Models
Maijunxian Wang, Yijiang Li, Bingyang Wang +6
The paper introduces FlipSet, a benchmark that tests vision‑language models on Level‑2 visual perspective taking by requiring them to mentally rotate 2D character strings, and find…
Vision Language Models Cannot Reason About Physical Transformation
Dezhi Luo, Yijiang Li, Maijunxian Wang +7
Understanding physical transformations is fundamental for reasoning in dynamic environments. While Vision Language Models (VLMs) show promise in embodied applications, whether they…
Vision-Language Models Mistake Head Orientation for Gaze Direction: Nonverbal Conversation Cues
Zory Zhang, Pinyuan Feng, Bingyang Wang +7
Where someone looks is a nonverbal communication cue that children and adults readily use. How well can Vision-Language Models (VLMs) infer gaze targets? To construct evaluation st…
Increasing Computation Resolves Conflicts in Vision Language Models
Bingyang Wang, Yijiang Li, Yitong Qiao +7
Cognitive control, the ability to coordinate competing information sources in pursuit of goals, is fundamental to intelligent behavior. We systematically investigate whether Vision…
Core Knowledge Deficits in Multi-Modal Language Models
Yijiang Li, Qingying Gao, Tianwei Zhao +8
While Multi-modal Large Language Models (MLLMs) demonstrate impressive abilities over high-level perception and reasoning, their robustness in the wild remains limited, often falli…
Achilles Heel of Distributed Multi-Agent Systems
Yiting Zhang, Yijiang Li, Tianwei Zhao +3
Multi-agent system (MAS) has demonstrated exceptional capabilities in addressing complex challenges, largely due to the integration of multiple large language models (LLMs). Howeve…