From the 1 of 5 linked papers with an AI index.
5 papers
GPOcc++: Unified Sparse Gaussian Occupancy Prediction with Visual Geometry Priors
Changqing Zhou, Yueru Luo, Yulan Guo +3
The paper introduces GPOcc and its extension GPOcc++, which turn visual geometry priors into sparse Gaussian occupancy representations for efficient 3D scene modeling, supporting b…
HTNav: A Hybrid Navigation Framework with Tiered Structure for Urban Aerial Vision-and-Language Navigation
Chengjie Fan, Cong Pan, Zijian Liu +2
Inspired by the general Vision-and-Language Navigation (VLN) task, aerial VLN has attracted widespread attention, owing to its significant practical value in applications such as l…
DecoVLN: Decoupling Observation, Reasoning, and Correction for Vision-and-Language Navigation
Zihao Xin, Wentong Li, Yixuan Jiang +4
Vision-and-Language Navigation (VLN) requires agents to follow long-horizon instructions and navigate complex 3D environments. However, existing approaches face two major challenge…
AgentVLN: Towards Agentic Vision-and-Language Navigation
Zihao Xin, Wentong Li, Yixuan Jiang +6
Vision-and-Language Navigation (VLN) requires an embodied agent to ground complex natural-language instructions into long-horizon navigation in unseen environments. While Vision-La…
MUVR: A Multi-Modal Untrimmed Video Retrieval Benchmark with Multi-Level Visual Correspondence
Yue Feng, Jinwei Hu, Qijia Lu +11
We propose the Multi-modal Untrimmed Video Retrieval task, along with a new benchmark (MUVR) to advance video retrieval for long-video platforms. MUVR aims to retrieve untrimmed vi…