5 papers
-0: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation
Zhe Li, Zhenzhe Zhang, Yangyang Wei +8
Humanoid household tasks often require concurrent loco-manipulation, where the robot must move, adjust posture, maintain balance, and manipulate objects as a single coordinated beh…
Action-to-Action Flow Matching
Jindou Jia, Gen Li, Xiangyu Chen +5
Diffusion-based policies have recently achieved remarkable success in robotics by formulating action prediction as a conditional denoising process. However, the standard practice o…
OTPrune: Distribution-Aligned Visual Token Pruning via Optimal Transport
Xiwen Chen, Wenhui Zhu, Gen Li +9
Multi-modal large language models (MLLMs) achieve strong visual-language reasoning but suffer from high inference cost due to redundant visual tokens. Recent work explores visual t…
RoboForge: Physically Optimized Text-guided Whole-Body Locomotion for Humanoids
Xichen Yuan, Zhe Li, Bofan Lyu +4
While generative models have become effective at producing human-like motions from text, transferring these motions to humanoid robots for physical execution remains challenging. E…
Trace-Focused Diffusion Policy for Multi-Modal Action Disambiguation in Long-Horizon Robotic Manipulation
Yuxuan Hu, Xiangyu Chen, Chuhao Zhou +4
Generative model-based policies have shown strong performance in imitation-based robotic manipulation by learning action distributions from demonstrations. However, in long-horizon…