13 papers
WNM-3D: A World Navigation Model with 3D Scene Conditioning for Closed-Loop VLN
Yuehao Huang, Yunzi Wu, Xiaotao Zhang +7
Recent vision-language navigation (VLN) systems increasingly adapt pretrained vision-language models (VLMs) into vision-language-action (VLA) policies that map egocentric observati…
TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging
Shengzhuo Yang, Ronghao Yu, Chuanjie Lv +5
Vision-language-action (VLA) models aim to understand natural-language instructions and visual observations, and to generate and execute corresponding actions as embodied agents. R…
SAGE-Nav: Leveraging LLM Planning and Alignment Fusion for Hierarchical Scene Graph-Guided Navigation
Hao Su, Yuehao Huang, Yukai Ma +2
Object-Goal Navigation (ObjNav) requires embodied agents to autonomously locate specified targets using only egocentric visual observations. Existing monolithic methods struggle wi…
DriveStack-VLA: Render-Teacher Alignment for BEV-Based DeepStack Vision-Language-Action Model
Jingke Wang, Zhenru Zhao, Shuangming Lei +8
Vision-Language-Action driving models convert a pretrained Vision-Language Model into a driving policy, allowing them to use world knowledge and follow language guidances. However,…
SparseWorld: Enhancing End-to-End Autonomous Driving via World Models with Sparse Scene Representation
Ruoyu Wang, Jingke Wang, Yukai Ma +5
Recently, world models have made significant progress in enhancing end-to-end driving systems through both future situation forecasting and improved scene understanding. However, e…
Monocular 3D Occupancy Perception for Robots on Sidewalks via Hybrid 2D-3D Learning
Yukai Ma, Joe Lin, Liu Liu +5
Sidewalks in the real world are crowded, cluttered, and less structured than roads, making 3D occupancy prediction a key ingredient for the safe navigation of mobile robots such as…