2 papers
cs.RO2026
-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation
Xiaowei Cai, Yunuo Cai, Bingao Chen +36
Long-horizon robot manipulation requires a robot to both execute individual skills reliably and sequence them coherently over extended tasks. Most hierarchical vision-language-acti…
cs.AI2026
PGN: Design and Implementation of a Vision-Language Navigation System Based on Pangu Multimodal Foundation Model
Li Xian, Mingxi Li, Yizheng Wang +3
Vision-Language Navigation (VLN) requires an embodied agent to interpret a natural-language instruction and predict actions from temporally ordered visual observations. Adapting a…