1 paper
Runze Xu, Yiluo Zhang, Jian Wang +2
Training generalist Vision-Language-Action(VLA) models typically requires massive, diverse robotic datasets with high-fidelity action annotations. While egocentric human manipulati…