From the 1 of 11 linked papers with an AI index.
11 papers
UniNav: A Unified World-Action Diffusion Model for Visual Navigation
Changqing Zhou, Yueru Luo, Zeyu Jiang +1
Image-goal visual navigation is a fundamental capability for embodied agents. Existing navigation policies efficiently predict waypoint trajectories but lack visual foresight, whil…
Reasoning to Regulate: Chain-of-Thought for Traffic Rule Understanding
Yueru Luo, Xu Yan, Changqing Zhou +5
Understanding and complying with traffic regulations is a safety-critical requirement for autonomous driving, yet remains challenging due to the diversity and context dependence of…
GPOcc++: Unified Sparse Gaussian Occupancy Prediction with Visual Geometry Priors
Changqing Zhou, Yueru Luo, Yulan Guo +3
The paper introduces GPOcc and its extension GPOcc++, which turn visual geometry priors into sparse Gaussian occupancy representations for efficient 3D scene modeling, supporting b…
Monocular Open Vocabulary Occupancy Prediction for Indoor Scenes
Changqing Zhou, Yueru Luo, Han Zhang +2
Open-vocabulary 3D occupancy is vital for embodied agents, which need to understand complex indoor environments where semantic categories are abundant and evolve beyond fixed taxon…
Reliev3R: Relieving Feed-forward Reconstruction from Multi-View Geometric Annotations
Youyu Chen, Junjun Jiang, Yueru Luo +4
With recent advances, Feed-forward Reconstruction Models (FFRMs) have demonstrated great potential in reconstruction quality and adaptiveness to multiple downstream tasks. However,…
WPT: World-to-Policy Transfer via Online World Model Distillation
Guangfeng Jiang, Yueru Luo, Jun Liu +6
Recent years have witnessed remarkable progress in world models, which primarily aim to capture the spatio-temporal correlations between an agent's actions and the evolving environ…