7 papers · 1 filter
Proxy Policy Steering
Chuanruo Ning, Tianrui Wang, Wei-Chiu Ma +1
Generalist robot policies carry broad manipulation priors from large-scale data, but specializing them to a new task remains the deployment bottleneck. This requires eliciting task…
Decoding Task Progress from VLA Representations
Atiksh Bhardwaj, Edward Weiyi Duan, Prithwish Dan +2
Vision-language-action models (VLAs) are moving rapidly towards deployment as general-purpose manipulation policies, but we currently lack basic tools for understanding what these…
X-Diffusion: Training Diffusion Policies on Cross-Embodiment Human Demonstrations
Maximus A. Pace, Prithwish Dan, Chuanruo Ning +5
Human videos are a scalable source of training data for robot learning. However, humans and robots significantly differ in embodiment, making many human actions infeasible for dire…
Implicit State Estimation via Video Replanning
Po-Chen Ko, Jiayuan Mao, Yu-Hsiang Fu +5
Video-based representations have gained prominence in planning and decision-making due to their ability to encode rich spatiotemporal dynamics and geometric relationships. These re…
Prompting with the Future: Open-World Model Predictive Control with Interactive Digital Twins
Chuanruo Ning, Kuan Fang, Wei-Chiu Ma
Recent advancements in open-world robot manipulation have been largely driven by vision-language models (VLMs). While these models exhibit strong generalization ability in high-lev…
X-Sim: Cross-Embodiment Learning via Real-to-Sim-to-Real
Prithwish Dan, Kushal Kedia, Angela Chao +4
Human videos offer a scalable way to train robot manipulation policies, but lack the action labels needed by standard imitation learning algorithms. Existing cross-embodiment appro…