1 paper
Nanxi Li, Xiang Wang, Yuanjie Chen +3
While Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in image and video understanding, their ability to comprehend the physical world has become…