1 paper
Wanyue Zhang, Wenxiang Wu, Wang Xu +6
Vision-language models (VLMs) have shown strong performance on static visual understanding, yet they still struggle with dynamic spatial reasoning that requires imagining how scene…