3 papers
cs.RO2026
HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model
Xiang Zhu, Puzhen Yuan, Yichen Liu +1
Learning generalizable vision-language-action (VLA) models from large-scale human videos is promising but challenging due to cross-embodiment discrepancies in both visual observati…
cs.RO2026
Learning Generalizable Robot Policy with Human Demonstration Video as a Prompt
Xiang Zhu, Yichen Liu, Hezhong Li +1
Recent robot learning methods commonly rely on imitation learning from massive robotic dataset collected with teleoperation. When facing a new task, such methods generally require…
cs.CV2025
UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
Jianke Zhang, Yanjiang Guo, Yucheng Hu +3
Recent advancements in Vision-Language-Action (VLA) models have leveraged pre-trained Vision-Language Models (VLMs) to improve the generalization capabilities. VLMs, typically pre-…