1 paper
Zuojin Tang, Shengchao Yuan, Xiaoxin Bai +4
Vision-language-action (VLA) models increasingly rely on auxiliary world modules to plan over long horizons, yet how such modules should be parameterized on top of a pretrained VLA…