1 paper
Chengye Song, Jiawei Zhang, Rui Song +5
Post-training of vision-language-action (VLA) models typically relies on expert demonstrations and policy interaction trajectories. However, in long-horizon manipulation, a single…