7 papers
DeVA: Decoupled Video-Action Model with physical guidance for robot policy learning
Mengqi Zhang, Sahil Khose, Simar Kareer +3
Generalizable robot manipulation requires policies that can anticipate how visual scenes evolve while executing language instructions. While recent Vision-Language-Action models be…
EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World
Ryan Punamiya, Simar Kareer, Zeyi Liu +37
Robot learning increasingly depends on large and diverse data, yet robot data collection remains expensive and difficult to scale. Egocentric human data offer a promising alternati…
Posterior Augmented Flow Matching
George Stoica, Sayak Paul, Matthew Wallingford +6
Flow matching (FM) trains a time-dependent vector field that transports samples from a simple prior to a complex data distribution. However, for high-dimensional images, each train…
Resolving Interference (RI): Disentangling Models for Improved Model Merging
Pratik Ramesh, George Stoica, Arun Iyer +2
Model merging has shown that multitask models can be created by directly combining the parameters of different models that are each specialized on tasks of interest. However, model…
Emergence of Human to Robot Transfer in Vision-Language-Action Models
Simar Kareer, Karl Pertsch, James Darpinian +5
Vision-language-action (VLA) models can enable broad open world generalization, but require large and diverse datasets. It is appealing to consider whether some of this data can co…
EgoBridge: Domain Adaptation for Generalizable Imitation from Egocentric Human Data
Ryan Punamiya, Dhruv Patel, Patcharapong Aphiwetsa +5
Egocentric human experience data presents a vast resource for scaling up end-to-end imitation learning for robotic manipulation. However, significant domain gaps in visual appearan…