11 citations · 109 across the 37 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2022★ 1 cited
LiteVL: Efficient Video-Language Learning with Enhanced Spatial-Temporal Modeling
Dongsheng Chen, Chaofan Tao, Lu Hou +3
Recent large-scale video-language pre-trained models have shown appealing performance on various downstream tasks. However, the pre-training process is computationally expensive du…
cs.CV2022
UTC: A Unified Transformer with Inter-Task Contrastive Learning for Visual Dialog
Cheng Chen, Yudong Zhu, Zhenshan Tan +4
Visual Dialog aims to answer multi-round, interactive questions based on the dialog history and image content. Existing methods either consider answer ranking and generating indivi…