17 papers
RoboTTT: Context Scaling for Robot Policies
Yunfan Jiang, Yevgen Chebotar, Ruijie Zheng +8
The paper introduces RoboTTT, a robot policy that uses test-time training to handle up to 8,000 timesteps of visual‑motor context, enabling one‑shot imitation from video, on‑the‑fl…
SimFoundry: Modular and Automated Scene Generation for Policy Learning and Evaluation
Nadun Ranawaka, Josiah Wong, Wei-Lin Pai +15
Training and evaluating robot policies in the real world is costly and difficult to scale. We introduce SimFoundry, a modular and automated system for zero-shot real-to-sim scene c…
Cosmos 3: Omnimodal World Models for Physical AI
NVIDIA, :, Aditi +293
We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-t…
Contrastive Action-Image Pre-training for Visuomotor Control
Yuvan Sharma, Dantong Niu, Anirudh Pai +16
Existing vision encoders for robotics face a fundamental bottleneck: robotic datasets lack the scale necessary for large-scale pre-training. Prior work circumvents this data scarci…
GRAIL: Generating Humanoid Loco-Manipulation from 3D Assets and Video Priors
Tianyi Xie, Haotian Zhang, Jinhyung Park +17
Scaling humanoid loco-manipulation requires robot-compatible demonstrations across diverse objects, whole-body motions, and scene geometries, but teleoperation and motion capture a…
HumanoidMimicGen: Data Generation for Loco-Manipulation via Whole-Body Planning
Kevin Lin, Ajay Mandlekar, Caelan Reed Garrett +7
Imitation learning is a promising approach for training humanoid robots to both walk and manipulate, but it requires a large number of demonstrations, which are time-intensive and…