7 papers · 1 filter
LD4WAM: Learning Latent Dynamics from Human Videos for World Action Models
Zhenhao Shen, Jiaqi Liang, Jasper Lu +11
Human video is playing an increasingly central role in training World Action Models (WAMs), owing to its diversity and low collection cost relative to teleoperated robot data. Howe…
Guided Streaming Stochastic Interpolant Policy
Puming Jiang, Meiyi Wang, Kelvin Lin +2
Inference-time guidance is essential for steering generative robot policies toward dynamic objectives without retraining, yet existing methods are largely confined to chunk-based a…
ForceVLA2: Unleashing Hybrid Force-Position Control with Force Awareness for Contact-Rich Manipulation
Yang Li, Zhaxizhuoma, Hongru Jiang +11
Embodied intelligence for contact-rich manipulation has predominantly relied on position control, while explicit awareness and regulation of interaction forces remain under-explore…
SkillVLA: Tackling Combinatorial Diversity in Dual-Arm Manipulation via Skill Reuse
Xuanran Zhai, Zekai Huang, Longyan Wu +5
Recent progress in vision-language-action (VLA) models has demonstrated strong potential for dual-arm manipulation, enabling complex behaviors and generalization to unseen environm…
CHD: Coupled Hierarchical Diffusion for Long-Horizon Tasks
Ce Hao, Anxing Xiao, Zhiwei Xue +1
Diffusion-based planners have shown strong performance in short-horizon tasks but often fail in complex, long-horizon settings. We trace the failure to loose coupling between high-…
DISCO: Language-Guided Manipulation with Diffusion Policies and Constrained Inpainting
Ce Hao, Kelvin Lin, Zhiwei Xue +2
Diffusion policies have demonstrated strong performance in generative modeling, making them promising for robotic manipulation guided by natural language instructions. However, gen…