2 papers
cs.CV2022
Dual-Level Decoupled Transformer for Video Captioning
Yiqi Gao, Xinglin Hou, Wei Suo +4
Video captioning aims to understand the spatio-temporal semantic concept of the video and generate descriptive sentences. The de-facto approach to this task dictates a text generat…
cs.CV2021
Proposal-free One-stage Referring Expression via Grid-Word Cross-Attention
Wei Suo, Mengyang Sun, Peng Wang +1
Referring Expression Comprehension (REC) has become one of the most important tasks in visual reasoning, since it is an essential step for many vision-and-language tasks such as vi…