1 paper
Zhihui Yang, Yupei Wang, Kaijie Mo +2
Despite significant progress in multimodal language models (LMs), it remains unclear whether visual grounding enhances their understanding of embodied knowledge compared to text-on…