3 papers
cs.CV2026
OTPrune: Distribution-Aligned Visual Token Pruning via Optimal Transport
Xiwen Chen, Wenhui Zhu, Gen Li +9
Multi-modal large language models (MLLMs) achieve strong visual-language reasoning but suffer from high inference cost due to redundant visual tokens. Recent work explores visual t…
cs.RO2026
RoboForge: Physically Optimized Text-guided Whole-Body Locomotion for Humanoids
Xichen Yuan, Zhe Li, Bofan Lyu +4
While generative models have become effective at producing human-like motions from text, transferring these motions to humanoid robots for physical execution remains challenging. E…
cs.RO2026
Trace-Focused Diffusion Policy for Multi-Modal Action Disambiguation in Long-Horizon Robotic Manipulation
Yuxuan Hu, Xiangyu Chen, Chuhao Zhou +4
Generative model-based policies have shown strong performance in imitation-based robotic manipulation by learning action distributions from demonstrations. However, in long-horizon…