4 papers
HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model
Xiang Zhu, Puzhen Yuan, Yichen Liu +1
Learning generalizable vision-language-action (VLA) models from large-scale human videos is promising but challenging due to cross-embodiment discrepancies in both visual observati…
Learning Generalizable Robot Policy with Human Demonstration Video as a Prompt
Xiang Zhu, Yichen Liu, Hezhong Li +1
Recent robot learning methods commonly rely on imitation learning from massive robotic dataset collected with teleoperation. When facing a new task, such methods generally require…
Dexora: Open-source VLA for High-DoF Bimanual Dexterity
Zongzheng Zhang, Jingrui Pang, Zhuo Yang +22
Vision-Language-Action (VLA) models have recently become a central direction in embodied AI, but current systems are restricted to either dual-gripper control or single-arm dextero…
Prompt a Robot to Walk with Large Language Models
Yen-Jen Wang, Bike Zhang, Jianyu Chen +1
Large language models (LLMs) pre-trained on vast internet-scale data have showcased remarkable capabilities across diverse domains. Recently, there has been escalating interest in…