2 citations · 3 across the 14 of their papers we have counts for
14 papers · 1 filter
Monkey See, Can Monkey Do? A Benchmark for Evaluating Robot Skill Learning by Observation
Weiwei Gu, Anmol Gupta, Anant Sah +5
Learning from Observation (LfO) is a fundamental robotic capability that replicates how humans and animals socially learn from each other. Beyond its biological parallels, this mod…
Meanings and Measurements: Multi-Agent Probabilistic Grounding for Vision-Language Navigation
Swagat Padhan, Lakshya Jain, Bhavya Minesh Shah +3
Robots collaborating with humans must convert natural language goals into actionable, physically grounded decisions. For example, executing a command such as "go two meters to the…
StageCraft: Execution Aware Mitigation of Distractor and Obstruction Failures in VLA Models
Kartikay Milind Pangaonkar, Prabin Rath, Omkar Patil +1
Large scale pre-training on text and image data along with diverse robot demonstrations has helped Vision Language Action models (VLAs) to generalize to novel tasks, objects and sc…
You've Got a Golden Ticket: Improving Generative Robot Policies With A Single Noise Vector
Omkar Patil, Ondrej Biza, Thomas Weng +9
What happens when a pretrained generative robot policy is provided a constant initial noise as input, rather than repeatedly sampling it from a Gaussian? We demonstrate that the pe…
PokeNet: Learning Kinematic Models of Articulated Objects from Human Observations
Anmol Gupta, Weiwei Gu, Omkar Patil +2
Articulation modeling enables robots to learn joint parameters of articulated objects for effective manipulation which can then be used downstream for skill learning or planning. E…
Factorizing Diffusion Policies for Observation Modality Prioritization
Omkar Patil, Prabin Rath, Kartikay Pangaonkar +2
Diffusion models have been extensively leveraged for learning robot skills from demonstrations. These policies are conditioned on several observational modalities such as proprioce…