8 papers
Scalable Behavior Cloning with Open Data, Training, and Evaluation
Arthur Allshire, Himanshu Gaurav Singh, Ritvik Singh +15
We introduce ABC, a fully open-source stack for manipulation with behavior cloning. At its core is ABC-130K: the largest open-source teleoperation dataset to date, featuring 3,500…
Do as I Do: Dexterous Manipulation Data from Everyday Human Videos
Bhawna Paliwal, Haritheja Etukuru, William Liang +3
How can we scalably generate data for robotic manipulation, especially on human-like platforms such as dexterous multi-fingered hands? Learning from human videos has recently emerg…
DiPOD: Diffusion Policy Optimization without Drifting Apart
Haozhe Jiang, Haiwen Feng, Pieter Abbeel +3
RL post-training has become increasingly pivotal for improving diffusion policies, but existing diffusion policy-gradient methods are often unstable and cannot achieve reliable pol…
SARM2: Multi-Task Stage Aware Reward Modeling for Self Improving Robotic Manipulation
Qianzhong Chen, Hau Zheng, Justin Yu +8
Fine-tuning vision-language-action (VLA) policies for long-horizon manipulation still relies heavily on behavior cloning, which requires costly high-quality demonstrations and keep…
Reward-Conditioned Reinforcement Learning
Michal Nauman, Marek Cygan, Pieter Abbeel
Single-task RL agents are typically trained under a fixed reward function, which limits their robustness to reward misspecification and their ability to adapt to changing preferenc…
When Does Non-Uniform Replay Matter in Reinforcement Learning?
Michal Korniak, MikoÅaj Czarnecki, Yarden As +3
Modern off-policy reinforcement learning algorithms often rely on simple uniform replay sampling and it remains unclear when and why non-uniform replay improves over this strong ba…