4 papers
LAP: Language-Action Pre-Training Enables Zero-shot Cross-Embodiment Transfer
Lihan Zha, Asher J. Hancock, Mingtong Zhang +5
A long-standing goal in robotics is a generalist policy that can be deployed zero-shot on new robot embodiments without per-embodiment adaptation. Despite large-scale multi-embodim…
Actions as Language: Fine-Tuning VLMs into VLAs Without Catastrophic Forgetting
Asher J. Hancock, Xindi Wu, Lihan Zha +2
Fine-tuning vision-language models (VLMs) on robot teleoperation data to create vision-language-action (VLA) models is a promising paradigm for training generalist policies, but it…
Is Your Imitation Learning Policy Better than Mine? Policy Comparison with Near-Optimal Stopping
David Snyder, Asher James Hancock, Apurva Badithela +6
Imitation learning has enabled robots to perform complex, long-horizon tasks in challenging dexterous manipulation settings. As new methods are developed, they must be rigorously e…
Run-time Observation Interventions Make Vision-Language-Action Models More Visually Robust
Asher J. Hancock, Allen Z. Ren, Anirudha Majumdar
Vision-language-action (VLA) models trained on large-scale internet data and robot demonstrations have the potential to serve as generalist robot policies. However, despite their l…