2 papers
cs.CL2025
Visual Room 2.0: Seeing is Not Understanding for MLLMs
Haokun Li, Yazhou Zhang, Jizhi Ding +2
Can multi-modal large language models (MLLMs) truly understand what they can see? Extending Searle's Chinese Room into the multi-modal domain, this paper proposes the Visual Room a…
cs.CL2025
Seeing is Not Understanding: A Benchmark on Perception-Cognition Disparities in Large Language Models
Haokun Li, Yazhou Zhang, Jizhi Ding +2
With the rapid advancement of Multimodal Large Language Models (MLLMs), they have demonstrated exceptional capabilities across a variety of vision-language tasks. However, current…