collaborators

30 papers

cs.CV2026

SpatioLM: Towards General Physical Spatial Intelligence in Vision-Language Models

Jing Wu, Jianhua Wu, Jiayi Guan +5

Vision-Language Models (VLMs) perform well on commonsense reasoning tasks but struggle with visual spatial reasoning. Most existing solutions introduce extra 3D prior inputs or ext…

cs.RO2026

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

Xiaomi Robotics Team, Jun Guo, Piaopiao Jin +31

We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulatio…

cs.AI2026

Text-Driven 3D Indoor Scene Synthesis in Non-Manhattan Environments

Xianhui Meng, Zirui Song, Yuchen Zhang +10

Large Language Models (LLMs) have demonstrated remarkable capabilities in 3D indoor synthesis for Manhattan environments. However, existing methods often fail to capture plausible…

cs.RO2026

FutureNav: Unified World-Action Modeling for Vision-and-Language Navigation

Lingfeng Zhang, Zeying Gong, Xiaoshuai Hao +7

Vision-and-language navigation (VLN) in continuous environments requires an agent to ground instructions in egocentric observations while maintaining spatial understanding across l…

cs.CV2026

Resonant Minds: Closed-Loop Social Avatars with Theory of Mind

Jianxu Shangguan, Jing Xu, Hang Ye +4

Creating lifelike digital humans with genuine social intelligence requires unifying cognitive reasoning and multimodal generation within a coherent framework. Current approaches tr…

cs.CV2026

DriveReward: A Comprehensive Dataset and Generative Vision-Language Reward Model for Autonomous Driving

Qimao Chen, Fang Li, Yuechen Luo +11

Reward models play a pivotal role in reinforcement learning (RL) and multi-modal trajectory selection for autonomous driving. However, acquiring such rewards typically relies on ha…