6 citations · 6 across the 2 of their papers we have counts for
3 papers
cs.CV2021
CLIP4Caption: CLIP for Video Caption
Mingkang Tang, Zhanyu Wang, Zhenhua Liu +3
Video captioning is a challenging task since it requires generating sentences describing various diverse and complex videos. Existing video captioning models lack adequate visual r…
cs.CL2021★ 6 cited
LICHEE: Improving Language Model Pre-training with Multi-grained Tokenization
Weidong Guo, Mingjun Zhao, Lusheng Zhang +5
Language model pre-training based on large corpora has achieved tremendous success in terms of constructing enriched contextual representations and has led to significant performan…
cs.CV2020
Pre-Trained Image Processing Transformer
Hanting Chen, Yunhe Wang, Tianyu Guo +7
As the computing power of modern hardware is increasing strongly, pre-trained deep learning models (e.g., BERT, GPT-3) learned on large-scale datasets have shown their effectivenes…