6 papers
Flex-: A Multi-Stream World-Action Model with Compute Flexibility
Ge Yan, Jinghao Liu, Yuzhi Fan +4
World-action models (WAMs) predict the future to act better, but nearly all of them predict only RGB latents, trained purely for pixel reconstruction, with no explicit signal for t…
Humanoid Policy ~ Human Policy
Ri-Zhao Qiu, Shiqi Yang, Xuxin Cheng +12
Training manipulation policies for humanoid robots with diverse data enhances their robustness and generalization across tasks and platforms. However, learning solely from robot de…
ManiFlow: A General Robot Manipulation Policy via Consistency Flow Training
Ge Yan, Jiyue Zhu, Yuquan Deng +8
This paper introduces ManiFlow, a visuomotor imitation learning policy for general robot manipulation that generates precise, high-dimensional actions conditioned on diverse visual…
Towards Embodiment Scaling Laws in Robot Locomotion
Bo Ai, Liu Dai, Nico Bohlinger +7
Cross-embodiment generalization underpins the vision of building generalist embodied agents for any robot, yet its enabling factors remain poorly understood. We investigate embodim…
Open X-Embodiment: Robotic Learning Datasets and RT-X Models
Embodiment Collaboration, Abby O'Neill, Abdul Rehman +291
Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, thi…
Integrating LMM Planners and 3D Skill Policies for Generalizable Manipulation
Yuelei Li, Ge Yan, Annabella Macaluso +3
The recent advancements in visual reasoning capabilities of large multimodal models (LMMs) and the semantic enrichment of 3D feature fields have expanded the horizons of robotic ca…