2 papers
cs.CV2022
Visual Commonsense-aware Representation Network for Video Captioning
Pengpeng Zeng, Haonan Zhang, Lianli Gao +3
Generating consecutive descriptions for videos, i.e., Video Captioning, requires taking full advantage of visual representation along with the generation process. Existing video ca…
cs.CV2022
Support-set based Multi-modal Representation Enhancement for Video Captioning
Xiaoya Chen, Jingkuan Song, Pengpeng Zeng +2
Video captioning is a challenging task that necessitates a thorough comprehension of visual scenes. Existing methods follow a typical one-to-one mapping, which concentrates on a li…