10 papers
UniNav: A Unified World-Action Diffusion Model for Visual Navigation
Changqing Zhou, Yueru Luo, Zeyu Jiang +1
Image-goal visual navigation is a fundamental capability for embodied agents. Existing navigation policies efficiently predict waypoint trajectories but lack visual foresight, whil…
GPOcc++: Unified Sparse Gaussian Occupancy Prediction with Visual Geometry Priors
Changqing Zhou, Yueru Luo, Yulan Guo +3
The paper introduces GPOcc and its extension GPOcc++, which turn visual geometry priors into sparse Gaussian occupancy representations for efficient 3D scene modeling, supporting b…
AdvNav: Behavior-Guided Black-Box Adversarial Attacks on Vision-Language Navigation
Chenyang Li, Kaige Li, Zeyu Jiang +1
The paper introduces AdvNav, a gradient‑free black‑box adversarial attack that perturbs first‑person visual inputs to disrupt vision‑and‑language navigation agents, using behavior‑…
FlowDec: Temporal Conditional Flow Decorruptor for Robust Continuous Vision-Language Navigation
Yufei Zhang, Changhao Chen
Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires agents to follow natural-language instructions in unseen scenes. While Large Models (LMs) have advanced…
Improving Adversarial Transferability on Vision-Language Pre-training Models via Surrogate-Specific Bias Correction
Lijia Yu, Jiuxin Cao, Yuchen Qiang +3
Adversarial examples reveal vulnerabilities in Vision-Language Pre-training (VLP) models and provide insights for improving robustness. A key property is cross-model transferabilit…
Monocular Open Vocabulary Occupancy Prediction for Indoor Scenes
Changqing Zhou, Yueru Luo, Han Zhang +2
Open-vocabulary 3D occupancy is vital for embodied agents, which need to understand complex indoor environments where semantic categories are abundant and evolve beyond fixed taxon…