9 papers
Understanding and Mitigating the Video-Action Generalization Gap via Temporal Ratio
Utkarsh A. Mishra, Yongxin Chen, Danfei Xu +3
Generative video foundation models exhibit strong compositional priors, yet world-action models (WAMs) and video-action models (VAMs) often lose these priors after finetuning on ro…
GRAFT: Graph-Based Affordance Transfer via Part Correspondence
Mengying Lin, Utkarsh Mishra, Ajay Mandlekar +1
Generalizing robotic manipulation to unseen objects remains challenging, as learning-based approaches require many demonstrations and fail in few-shot settings. Prior work transfer…
Energy-based Compositional Diffusion Planning
Tao Sun, Utkarsh Aashu Mishra, Jiaxin Lu +2
Compositional diffusion planners aim to solve long-horizon robotic tasks using short training trajectories. Yet, current approaches often rely on the heuristic stitching of local p…
Coarse-to-Fine Compositional Diffusion for Long-Horizon Planning
Byoungwoo Park, Utkarsh A. Mishra, Jaemoo Choi +2
Diffusion models provide strong priors for generating structured data, but many tasks require outputs beyond the scale on which these models are typically trained. Compositional ge…
KinDER: A Physical Reasoning Benchmark for Robot Learning and Planning
Yixuan Huang, Bowen Li, Vaibhav Saxena +9
Robotic systems that interact with the physical world must reason about kinematic and dynamic constraints imposed by their own embodiment, their environment, and the task at hand.…
Compositional Visual Planning via Inference-Time Diffusion Scaling
Yixin Zhang, Yunhao Luo, Utkarsh Aashu Mishra +3
Diffusion models excel at short-horizon robot planning, yet scaling them to long-horizon tasks remains challenging due to computational constraints and limited training data. Exist…