Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
From Routes to Steps: Separating Semantic Progress from Local Execution in Vision-and-Language Navigation
Xiangyun Huang, Xiangchen Wang, Runfeng Lin +5
Vision-and-Language Navigation (VLN) requires an agent to follow a route-level instruction by executing its constituent steps from egocentric visual observations. Existing VLM-base…
cs.CV2026
\textsc{NaVIDA}: Vision-Language Navigation with Inverse Dynamics Augmentation
Weiye Zhu, Zekai Zhang, Xiangchen Wang +5
Vision-and-Language Navigation (VLN) requires agents to interpret natural language instructions and act coherently in visually rich environments. However, most existing methods rel…
cs.CV2025
Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors
Xiangchen Wang, Jinrui Zhang, Teng Wang +2
Recent advancements in large video-language models have revolutionized video understanding tasks. However, their efficiency is significantly constrained by processing high volumes…