8 papers
RoboSynChallenge: Mastering Real-World Dexterity via Generalizing Synthesized Manipulation Skills
Runyi Zhao, Ruixin Wu, Chengkun Li +15
Achieving generalizable robotic manipulation remains a central challenge in embodied intelligence. Despite rapid advances in model architectures and learning algorithms, progress i…
Vid2WAM: Distilling Video Diffusion Priors into World Action Models
Chenhao Qiu, Ruixiang Wang, Runyi Zhao +7
World Action Models (WAMs) improve robot policy learning by jointly modeling future visual dynamics and actions. However, their scalability and generalization remain constrained by…
LIRA: Local Cross-Layer Information Routing for Vision-Language-Action Decoding
Zhewei Zhang, Puyue Wang, Guanren Qiao +10
Vision-Language-Action (VLA) models transform representations from pretrained vision-language models (VLMs) into robot actions, yet the interface that routes intermediate VLM featu…
RoboReact: Agentic Skill Distillation from Generated Egocentric Videos for Generalizable Whole-Body Manipulation
Shuliang He, Shuai Wang, Bo Yue +3
Humanoid robots have the potential to perform dexterous manipulation in human environments, yet acquiring diverse and generalizable skills remains costly due to expensive hardware…
TacticGen: Grounding Adaptable and Scalable Generation of Football Tactics
Sheng Xu, Guiliang Liu, Tarak Kharrat +12
Success in association football relies on both individual skill and coordinated tactics. While recent advancements in spatio-temporal data and deep learning have enabled predictive…
DexWorldModel: Causal Latent World Modeling towards Automated Learning of Embodied Tasks
Yueci Deng, Guiliang Liu, Kui Jia
Deploying generative World-Action Models for manipulation is severely bottlenecked by redundant pixel-level reconstruction, memory scaling, and sequential inferenc…