1 paper
Yijun Yang, Shenghe Zheng, Wenbo Li +8
Despite the remarkable prowess of Vision-Language Models (VLMs) in general multimodal tasks, they remain fundamentally ``flat'' when reasoning about the physical world. We argue th…