collaborators

9 papers

cs.CV2026

Understanding and Mitigating the Video-Action Generalization Gap via Temporal Ratio

Utkarsh A. Mishra, Yongxin Chen, Danfei Xu +3

Generative video foundation models exhibit strong compositional priors, yet world-action models (WAMs) and video-action models (VAMs) often lose these priors after finetuning on ro…

cs.RO2026

GRAFT: Graph-Based Affordance Transfer via Part Correspondence

Mengying Lin, Utkarsh Mishra, Ajay Mandlekar +1

Generalizing robotic manipulation to unseen objects remains challenging, as learning-based approaches require many demonstrations and fail in few-shot settings. Prior work transfer…

cs.RO2026

Energy-based Compositional Diffusion Planning

Tao Sun, Utkarsh Aashu Mishra, Jiaxin Lu +2

Compositional diffusion planners aim to solve long-horizon robotic tasks using short training trajectories. Yet, current approaches often rely on the heuristic stitching of local p…

cs.RO2026

Coarse-to-Fine Compositional Diffusion for Long-Horizon Planning

Byoungwoo Park, Utkarsh A. Mishra, Jaemoo Choi +2

Diffusion models provide strong priors for generating structured data, but many tasks require outputs beyond the scale on which these models are typically trained. Compositional ge…

cs.RO2026

KinDER: A Physical Reasoning Benchmark for Robot Learning and Planning

Yixuan Huang, Bowen Li, Vaibhav Saxena +9

Robotic systems that interact with the physical world must reason about kinematic and dynamic constraints imposed by their own embodiment, their environment, and the task at hand.…

cs.RO2026

Compositional Visual Planning via Inference-Time Diffusion Scaling

Yixin Zhang, Yunhao Luo, Utkarsh Aashu Mishra +3

Diffusion models excel at short-horizon robot planning, yet scaling them to long-horizon tasks remains challenging due to computational constraints and limited training data. Exist…