30 papers
SpatioLM: Towards General Physical Spatial Intelligence in Vision-Language Models
Jing Wu, Jianhua Wu, Jiayi Guan +5
Vision-Language Models (VLMs) perform well on commonsense reasoning tasks but struggle with visual spatial reasoning. Most existing solutions introduce extra 3D prior inputs or ext…
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories
Xiaomi Robotics Team, Jun Guo, Piaopiao Jin +31
We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulatio…
Text-Driven 3D Indoor Scene Synthesis in Non-Manhattan Environments
Xianhui Meng, Zirui Song, Yuchen Zhang +10
Large Language Models (LLMs) have demonstrated remarkable capabilities in 3D indoor synthesis for Manhattan environments. However, existing methods often fail to capture plausible…
FutureNav: Unified World-Action Modeling for Vision-and-Language Navigation
Lingfeng Zhang, Zeying Gong, Xiaoshuai Hao +7
Vision-and-language navigation (VLN) in continuous environments requires an agent to ground instructions in egocentric observations while maintaining spatial understanding across l…
Resonant Minds: Closed-Loop Social Avatars with Theory of Mind
Jianxu Shangguan, Jing Xu, Hang Ye +4
Creating lifelike digital humans with genuine social intelligence requires unifying cognitive reasoning and multimodal generation within a coherent framework. Current approaches tr…
DriveReward: A Comprehensive Dataset and Generative Vision-Language Reward Model for Autonomous Driving
Qimao Chen, Fang Li, Yuechen Luo +11
Reward models play a pivotal role in reinforcement learning (RL) and multi-modal trajectory selection for autonomous driving. However, acquiring such rewards typically relies on ha…