6 papers
SeededGrasp: Language-Guided Grasping in Complex Scenes with Multiple Embodiments
Yang Xu, Gurpreet Singh Mukker, Raymond Wang +3
Practical robotic grasping in complex scenes requires both 3D spatial reasoning and alignment with task-specific requirements. Vision-language models (VLMs) offer a natural way to…
Update-Free On-Policy Steering via Verifiers
Maria Attarian, Ian Vyse, Claas Voelcker +7
In recent years, Behavior Cloning (BC) has become one of the most prevalent methods for learning manipulation from human demonstrations. Despite their successes, BC policies are of…
BPP: Long-Context Robot Imitation Learning by Focusing on Key History Frames
Max Sobol Mark, Jacky Liang, Maria Attarian +4
Many robot tasks require attending to the history of past observations. For example, finding an item in a room requires remembering which places have already been searched. However…
GeoMatch++: Morphology Conditioned Geometry Matching for Multi-Embodiment Grasping
Yunze Wei, Maria Attarian, Igor Gilitschenski
Despite recent progress on multi-finger dexterous grasping, current methods focus on single grippers and unseen objects, and even the ones that explore cross-embodiment, often fail…
Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers
Vidhi Jain, Maria Attarian, Nikhil J Joshi +10
Large-scale multi-task robotic manipulation systems often rely on text to specify the task. In this work, we explore whether a robot can learn by observing humans. To do so, the ro…
Learning to Learn Faster from Human Feedback with Language Model Predictive Control
Jacky Liang, Fei Xia, Wenhao Yu +47
Large language models (LLMs) have been shown to exhibit a wide range of capabilities, such as writing robot code from language commands -- enabling non-experts to direct robot beha…