52 citations · 113 across the 26 of their papers we have counts for
Showing 2022Show all
3 papers · 1 filter
cs.CV2022
PIC 4th Challenge: Semantic-Assisted Multi-Feature Encoding and Multi-Head Decoding for Dense Video Captioning
Yifan Lu, Ziqi Zhang, Yuxin Chen +3
The task of Dense Video Captioning (DVC) aims to generate captions with timestamps for multiple events in one video. Semantic information plays an important role for both localizat…
cs.CV2022★ 11 cited
Improving Visual Grounding with Visual-Linguistic Verification and Iterative Reasoning
Li Yang, Yan Xu, Chunfeng Yuan +3
Visual grounding is a task to locate the target indicated by a natural language expression. Existing methods extend the generic object detection framework to this problem. They bas…
cs.CV2022★ 5 cited
CREATE: A Benchmark for Chinese Short Video Retrieval and Title Generation
Ziqi Zhang, Yuxin Chen, Zongyang Ma +5
Previous works of video captioning aim to objectively describe the video's actual content, which lacks subjective and attractive expression, limiting its practical application scen…