7 papers
World2VLM: Distilling World Model Imagination into VLMs for Dynamic Spatial Reasoning
Wanyue Zhang, Wenxiang Wu, Wang Xu +6
Vision-language models (VLMs) have shown strong performance on static visual understanding, yet they still struggle with dynamic spatial reasoning that requires imagining how scene…
Video2Layout: Recall and Reconstruct Metric-Grounded Cognitive Map for Spatial Reasoning
Yibin Huang, Wang Xu, Wanyue Zhang +6
Spatial intelligence is a critical frontier for Multimodal Large Language Models (MLLMs), empowering them to comprehend the physical world. Drawing inspiration from human perceptio…
MM-UAVBench: How Well Do Multimodal Large Language Models See, Think, and Plan in Low-Altitude UAV Scenarios?
Shiqi Dai, Zizhi Ma, Zhicong Luo +8
While Multimodal Large Language Models (MLLMs) have exhibited remarkable general intelligence across diverse domains, their potential in low-altitude applications dominated by Unma…
Scaling Spatial Reasoning in MLLMs through Programmatic Data Synthesis
Zhi Helu, Huang Jingjing, Xu Wang +9
Embodied intelligence, a grand challenge in artificial intelligence, is fundamentally constrained by the limited spatial understanding and reasoning capabilities of current models.…
Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture
Wanyue Zhang, Yibin Huang, Yangbin Xu +5
Spatial understanding is essential for Multimodal Large Language Models (MLLMs) to support perception, reasoning, and planning in embodied environments. Despite recent progress, ex…
CookBench: A Long-Horizon Embodied Planning Benchmark for Complex Cooking Scenarios
Muzhen Cai, Xiubo Chen, Yining An +5
Embodied Planning is dedicated to the goal of creating agents capable of executing long-horizon tasks in complex physical worlds. However, existing embodied planning benchmarks fre…