1 paper
Yu Chen, Ting Lei, Yaoyi Li +3
Multimodal large language models (MLLMs) combine linguistic reasoning with visual perception, yet their ability to perform visual spatial planning under explicit or previously unse…