2 papers
cs.CV2025
SE-VLN: A Self-Evolving Vision-Language Navigation Framework Based on Multimodal Large Language Models
Xiangyu Dong, Haoran Zhao, Jiang Gao +5
Recent advances in vision-language navigation (VLN) were mainly attributed to emerging large language models (LLMs). These methods exhibited excellent generalization capabilities i…
cs.CV2025
Global2Local: A Joint-Hierarchical Attention for Video Captioning
Chengpeng Dai, Fuhai Chen, Xiaoshuai Sun +3
Recently, automatic video captioning has attracted increasing attention, where the core challenge lies in capturing the key semantic items, like objects and actions as well as thei…