5 papers
DeVA: Decoupled Video-Action Model with physical guidance for robot policy learning
Mengqi Zhang, Sahil Khose, Simar Kareer +3
Generalizable robot manipulation requires policies that can anticipate how visual scenes evolve while executing language instructions. While recent Vision-Language-Action models be…
EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World
Ryan Punamiya, Simar Kareer, Zeyi Liu +37
Robot learning increasingly depends on large and diverse data, yet robot data collection remains expensive and difficult to scale. Egocentric human data offer a promising alternati…
Emergence of Human to Robot Transfer in Vision-Language-Action Models
Simar Kareer, Karl Pertsch, James Darpinian +5
Vision-language-action (VLA) models can enable broad open world generalization, but require large and diverse datasets. It is appealing to consider whether some of this data can co…
EMMA: Scaling Mobile Manipulation via Egocentric Human Data
Lawrence Y. Zhu, Pranav Kuppili, Ryan Punamiya +5
Scaling mobile manipulation imitation learning is bottlenecked by expensive mobile robot teleoperation. We present Egocentric Mobile MAnipulation (EMMA), an end-to-end framework tr…
EgoBridge: Domain Adaptation for Generalizable Imitation from Egocentric Human Data
Ryan Punamiya, Dhruv Patel, Patcharapong Aphiwetsa +5
Egocentric human experience data presents a vast resource for scaling up end-to-end imitation learning for robotic manipulation. However, significant domain gaps in visual appearan…