8 papers
Patch Policy: Efficient Embodied Control via Dense Visual Representations
Gaoyue Zhou, Zichen Jeff Cui, Ada Langford +3
Pretrained dense visual features from Vision Transformers (ViTs) are powerful yet have been underutilized in robot learning. Modern robot policies either compress each observation…
Temporal Straightening for Latent Planning
Ying Wang, Oumayma Bounou, Gaoyue Zhou +4
Learning good representations is essential for latent planning with world models. While pretrained visual encoders produce strong semantic visual features, they are not tailored to…
World Models for Learning Dexterous Hand-Object Interactions from Human Videos
Raktim Gautam Goswami, Amir Bar, David Fan +6
Modeling dexterous hand-object interactions is challenging as it requires understanding how subtle finger motions influence the environment through contact with objects. While rece…
Beyond Language Modeling: An Exploration of Multimodal Pretraining
Shengbang Tong, David Fan, John Nguyen +18
The visual world offers a critical axis for advancing foundation models beyond language. Despite growing interest in this direction, the design space for native multimodal models r…
Towards Embodiment Scaling Laws in Robot Locomotion
Bo Ai, Liu Dai, Nico Bohlinger +7
Cross-embodiment generalization underpins the vision of building generalist embodied agents for any robot, yet its enabling factors remain poorly understood. We investigate embodim…
Open X-Embodiment: Robotic Learning Datasets and RT-X Models
Embodiment Collaboration, Abby O'Neill, Abdul Rehman +291
Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, thi…