works on

From the 1 of 5 linked papers with an AI index.

collaborators

5 papers

cs.CV2026

GPOcc++: Unified Sparse Gaussian Occupancy Prediction with Visual Geometry Priors

Changqing Zhou, Yueru Luo, Yulan Guo +3

The paper introduces GPOcc and its extension GPOcc++, which turn visual geometry priors into sparse Gaussian occupancy representations for efficient 3D scene modeling, supporting b…

cs.RO2026

HTNav: A Hybrid Navigation Framework with Tiered Structure for Urban Aerial Vision-and-Language Navigation

Chengjie Fan, Cong Pan, Zijian Liu +2

Inspired by the general Vision-and-Language Navigation (VLN) task, aerial VLN has attracted widespread attention, owing to its significant practical value in applications such as l…

cs.RO2026

DecoVLN: Decoupling Observation, Reasoning, and Correction for Vision-and-Language Navigation

Zihao Xin, Wentong Li, Yixuan Jiang +4

Vision-and-Language Navigation (VLN) requires agents to follow long-horizon instructions and navigate complex 3D environments. However, existing approaches face two major challenge…

cs.RO2026

AgentVLN: Towards Agentic Vision-and-Language Navigation

Zihao Xin, Wentong Li, Yixuan Jiang +6

Vision-and-Language Navigation (VLN) requires an embodied agent to ground complex natural-language instructions into long-horizon navigation in unseen environments. While Vision-La…

cs.CV2025

MUVR: A Multi-Modal Untrimmed Video Retrieval Benchmark with Multi-Level Visual Correspondence

Yue Feng, Jinwei Hu, Qijia Lu +11

We propose the Multi-modal Untrimmed Video Retrieval task, along with a new benchmark (MUVR) to advance video retrieval for long-video platforms. MUVR aims to retrieve untrimmed vi…