2 papers
cs.CV2026
Efficient-VLN: A Simple yet Strong Baseline for Efficient Vision-Language Navigation
Duo Zheng, Shijia Huang, Yanyang Li +1
While Multimodal Large Language Models (MLLMs) have demonstrated significant promise in Vision-Language Navigation (VLN), existing agents remain heavily constrained by systemic bot…
cs.CV2026
GenSpan: Generation-Calibrated Motion Span Priors for Multi-Verb Video Corpus Moment Retrieval
Yunzhuo Sun, Xinyue Liu, Yanyang Li +4
Video Corpus Moment Retrieval (VCMR) aims to retrieve both the correct video and its temporal segment corresponding to a natural-language query, a task that is especially challengi…