1 paper · 1 filter
Yalcin Tur, Jalal Naghiyev, Haoquan Fang +4
Current Vision-Language-Action (VLA) models rely on fixed computational depth, expending the same amount of compute on simple adjustments and complex multi-step manipulation. While…