3 papers
cs.CV2026
DynaWM: A Base-VLA-Guided World Foundation Model for Moving-Object Manipulation
Chongkei Chang, Zhidong Deng
Although vision-language-action (VLA) models have received widespread attention, many challenges remain in manipulating dynamic moving objects. In most existing approaches, end-to-…
cs.RO2026
Key-Gram: Extensible World Knowledge for Embodied Manipulation
Jingjing Fan, Siyuan Li, Botao Ren +1
Embodied control increasingly requires models to follow compositional language instructions while reasoning over dynamic visual states. However, current vision-language-action poli…
cs.RO2026
FUTURE-VLA: Forecasting Unified Trajectories Under Real-time Execution
Jingjing Fan, Yushan Liu, Shoujie Li +5
General vision-language models increasingly support unified spatiotemporal reasoning over long video streams, yet deploying such capabilities on robots remains constrained by the p…