5 papers
LD4WAM: Learning Latent Dynamics from Human Videos for World Action Models
Zhenhao Shen, Jiaqi Liang, Jasper Lu +11
Human video is playing an increasingly central role in training World Action Models (WAMs), owing to its diversity and low collection cost relative to teleoperated robot data. Howe…
Towards Trustworthy Embodied Intelligence: A Systems Framework and Graded Trustworthiness Levels
Xinyu Yang, Tianxing Chen, Honghao Su +38
Embodied intelligence integrates learned perception and decision making with real-time computation, control, and physical interaction. Because failures can cause immediate physical…
AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding
Qize Yu, Jiadi You, Yuran Wang +10
Vision-Language-Action (VLA) models leverage the rich world knowledge of pretrained vision-language models (VLMs) to enable instruction-following robotic manipulation. However, the…
LongAct: Harnessing Intrinsic Activation Patterns for Long-Context Reinforcement Learning
Bowen Ping, Zijun Chen, Tingfeng Hui +4
Reinforcement Learning (RL) has emerged as a critical driver for enhancing the reasoning capabilities of Large Language Models (LLMs). While recent advancements have focused on rew…
A3D: Adaptive Affordance Assembly with Dual-Arm Manipulation
Jiaqi Liang, Yue Chen, Qize Yu +4
Furniture assembly is a crucial yet challenging task for robots, requiring precise dual-arm coordination where one arm manipulates parts while the other provides collaborative supp…