1 paper · 1 filter
Shengye Dong, Haochen Niu, Hao Liu +3
Modern vision-language-action (VLA) policies predict a whole chunk of actions: one to two seconds of coordinated motion emitted in a single forward pass. Yet an action chunk is ess…