24 papers
WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time
Yusen Feng, Bingchen Han, Jiangran Lyu +13
Steering robot foundation models (RFMs) toward new task variants or user-preferred behaviors remains challenging, often requiring additional robot demonstrations, task-specific fin…
MM-Nav: Multi-View VLA Model for Robust Visual Navigation via Multi-Expert Learning
Tianyu Xu, Jiawei Chen, Jiazhao Zhang +5
Visual navigation policy is widely regarded as a promising direction, as it mimics humans by using egocentric visual observations for navigation. However, optical information of vi…
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision
Shengliang Deng, Mi Yan, Yixin Zheng +7
While Vision-Language-Action (VLA) models excel in generalist manipulation, they often lack fine-grained spatial awareness and show limited viewpoint robustness. This limitation la…
Cosmos 3: Omnimodal World Models for Physical AI
NVIDIA, :, Aditi +293
We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-t…
CLAR: Learning 3D Representations for Robotic Manipulation by Fusing Masked Reconstruction with Multi-Level Contrastive Alignment
Wenbo Cui, Chengyang Zhao, Yuhui Chen +4
The spatial information inherent in 3D point clouds is crucial for robotic manipulation. However, existing 3D pre-training methods face a fundamental trade-off: Masked Autoencoding…
AllDayNav: Lifelong Navigation via Real-World Reinforcement Learning
Hang Yin, Yinan Liang, Jiazhao Zhang +4
Lifelong embodied navigation in dynamic environments requires robots to form persistent scene understanding from fragmentary observations, which remains difficult for existing meth…