5 citations · 7 across the 7 of their papers we have counts for
7 papers
GALA: Geometry-Aware Latent Action Modeling for Vision-Language-Action Model Pretraining across Embodiments
Yichen Liu, Puzhen Yuan, Xiang Zhu +2
Learning large-scale vision-language-action (VLA) models from multi-embodiment datasets remains challenging due to heterogeneous action spaces across end effectors. Although latent…
HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model
Xiang Zhu, Puzhen Yuan, Yichen Liu +1
Learning generalizable vision-language-action (VLA) models from large-scale human videos is promising but challenging due to cross-embodiment discrepancies in both visual observati…
Learning Generalizable Robot Policy with Human Demonstration Video as a Prompt
Xiang Zhu, Yichen Liu, Hezhong Li +1
Recent robot learning methods commonly rely on imitation learning from massive robotic dataset collected with teleoperation. When facing a new task, such methods generally require…
UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
Jianke Zhang, Yanjiang Guo, Yucheng Hu +3
Recent advancements in Vision-Language-Action (VLA) models have leveraged pre-trained Vision-Language Models (VLMs) to improve the generalization capabilities. VLMs, typically pre-…
Advancing Humanoid Locomotion: Mastering Challenging Terrains with Denoising World Model Learning
Xinyang Gu, Yen-Jen Wang, Xiang Zhu +4
Humanoid robots, with their human-like skeletal structure, are especially suited for tasks in human-centric environments. However, this structure is accompanied by additional chall…
Stylized Table Tennis Robots Skill Learning with Incomplete Human Demonstrations
Xiang Zhu, Zixuan Chen, Jianyu Chen
In recent years, Reinforcement Learning (RL) is becoming a popular technique for training controllers for robots. However, for complex dynamic robot control tasks, RL-based method…