5 papers
Understanding and Mitigating the Video-Action Generalization Gap via Temporal Ratio
Utkarsh A. Mishra, Yongxin Chen, Danfei Xu +3
Generative video foundation models exhibit strong compositional priors, yet world-action models (WAMs) and video-action models (VAMs) often lose these priors after finetuning on ro…
REFINE-DP: Diffusion Policy Fine-tuning for Humanoid Loco-manipulation via Reinforcement Learning
Zhaoyuan Gu, Yipu Chen, Zimeng Chai +12
Humanoid loco-manipulation requires coordinated task-space motion planning with stable loco-manipulation command tracking under complex robot-environment dynamics and long-horizon…
Compositional Visual Planning via Inference-Time Diffusion Scaling
Yixin Zhang, Yunhao Luo, Utkarsh Aashu Mishra +3
Diffusion models excel at short-horizon robot planning, yet scaling them to long-horizon tasks remains challenging due to computational constraints and limited training data. Exist…
Compositional Diffusion with Guided Search for Long-Horizon Planning
Utkarsh A Mishra, David He, Yongxin Chen +1
Generative models have emerged as powerful tools for planning, with compositional approaches offering particular promise for modeling long-horizon task distributions by composing t…
Joint Model-based Model-free Diffusion for Planning with Constraints
Wonsuhk Jung, Utkarsh A. Mishra, Nadun Ranawaka Arachchige +3
Model-free diffusion planners have shown great promise for robot motion planning, but practical robotic systems often require combining them with model-based optimization modules t…