2 papers
cs.CV2025
Global2Local: A Joint-Hierarchical Attention for Video Captioning
Chengpeng Dai, Fuhai Chen, Xiaoshuai Sun +3
Recently, automatic video captioning has attracted increasing attention, where the core challenge lies in capturing the key semantic items, like objects and actions as well as thei…
cs.CV2024
Image Captioning via Dynamic Path Customization
Yiwei Ma, Jiayi Ji, Xiaoshuai Sun +4
This paper explores a novel dynamic network for vision and language tasks, where the inferring structure is customized on the fly for different inputs. Most previous state-of-the-a…