4 papers
MoTVLA: A Vision-Language-Action Model with Unified Fast-Slow Reasoning
Wenhui Huang, Changhe Chen, Han Qi +3
Integrating visual-language instructions into visuomotor policies is gaining momentum in robot learning for enhancing open-world generalization. Despite promising advances, existin…
Flexible Locomotion Learning with Diffusion Model Predictive Control
Runhan Huang, Haldun Balim, Heng Yang +1
Legged locomotion demands controllers that are both robust and adaptable, while remaining compatible with task and safety considerations. However, model-free reinforcement learning…
SPIE: Semantic and Structural Post-Training of Image Editing Diffusion Models with AI feedback
Elior Benarous, Yilun Du, Heng Yang
This paper presents SPIE: a novel approach for semantic and structural post-training of instruction-based image editing diffusion models, addressing key challenges in alignment wit…
MoE-Loco: Mixture of Experts for Multitask Locomotion
Runhan Huang, Shaoting Zhu, Yilun Du +1
We present MoE-Loco, a Mixture of Experts (MoE) framework for multitask locomotion for legged robots. Our method enables a single policy to handle diverse terrains, including bars,…