2 papers
cs.CV2026
DynaWM: A Base-VLA-Guided World Foundation Model for Moving-Object Manipulation
Chongkei Chang, Zhidong Deng
Although vision-language-action (VLA) models have received widespread attention, many challenges remain in manipulating dynamic moving objects. In most existing approaches, end-to-…
cs.RO2026
FUTURE-VLA: Forecasting Unified Trajectories Under Real-time Execution
Jingjing Fan, Yushan Liu, Shoujie Li +5
General vision-language models increasingly support unified spatiotemporal reasoning over long video streams, yet deploying such capabilities on robots remains constrained by the p…