6 papers · 1 filter
OpenWAM: An Open, Modular Exploration Towards Systematic World-Action Model Pretraining
Yuran Wang, Siqiao Huang, Mingleyang Li +21
World-Action Models inherit world knowledge from video-generative priors, and channel it into executable control signals through embodied experience. Existing systems, however, are…
LD4WAM: Learning Latent Dynamics from Human Videos for World Action Models
Zhenhao Shen, Jiaqi Liang, Jasper Lu +11
Human video is playing an increasingly central role in training World Action Models (WAMs), owing to its diversity and low collection cost relative to teleoperated robot data. Howe…
Towards Trustworthy Embodied Intelligence: A Systems Framework and Graded Trustworthiness Levels
Xinyu Yang, Tianxing Chen, Honghao Su +38
Embodied intelligence integrates learned perception and decision making with real-time computation, control, and physical interaction. Because failures can cause immediate physical…
AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding
Qize Yu, Jiadi You, Yuran Wang +10
Vision-Language-Action (VLA) models leverage the rich world knowledge of pretrained vision-language models (VLMs) to enable instruction-following robotic manipulation. However, the…
SUGAR: A Scalable Human-Video-Driven Generalizable Humanoid Loco-Manipulation Learning Framework
Tianshu Wu, Xiangqi Kong, Yue Chen +5
Building humanoid robots capable of generalizable whole-body loco-manipulation in the real world remains a fundamental challenge. Existing methods either rely on laborious task-spe…
A3D: Adaptive Affordance Assembly with Dual-Arm Manipulation
Jiaqi Liang, Yue Chen, Qize Yu +4
Furniture assembly is a crucial yet challenging task for robots, requiring precise dual-arm coordination where one arm manipulates parts while the other provides collaborative supp…