9 papers · 1 filter
MEM: Multi-Scale Embodied Memory for Vision Language Action Models
Marcel Torne, Karl Pertsch, Homer Walke +14
Conventionally, memory in end-to-end robotic learning involves inputting a sequence of past observations into the learned policy. However, in complex multi-stage real-world tasks,…
Towards Embodiment Scaling Laws in Robot Locomotion
Bo Ai, Liu Dai, Nico Bohlinger +7
Cross-embodiment generalization underpins the vision of building generalist embodied agents for any robot, yet its enabling factors remain poorly understood. We investigate embodim…
Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models
Lucy Xiaoyang Shi, Brian Ichter, Michael Equi +12
Generalist robots that can perform a range of different tasks in open-world settings must be able to not only reason about the steps needed to accomplish their goals, but also proc…
Open X-Embodiment: Robotic Learning Datasets and RT-X Models
Embodiment Collaboration, Abby O'Neill, Abdul Rehman +291
Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, thi…
FAST: Efficient Action Tokenization for Vision-Language-Action Models
Karl Pertsch, Kyle Stachowicz, Brian Ichter +6
Autoregressive sequence models, such as Transformer-based vision-language action (VLA) policies, can be tremendously effective for capturing complex and generalizable robotic behav…
LLaMAR: Long-Horizon Planning for Multi-Agent Robots in Partially Observable Environments
Siddharth Nayak, Adelmo Morrison Orozco, Marina Ten Have +10
The ability of Language Models (LMs) to understand natural language makes them a powerful tool for parsing human instructions into task plans for autonomous robots. Unlike traditio…