12 citations · 23 across the 3 of their papers we have counts for
3 papers
cs.CV2021★ 9 cited
Towards a Unified Foundation Model: Jointly Pre-Training Transformers on Unpaired Images and Text
Qing Li, Boqing Gong, Yin Cui +4
In this paper, we explore the possibility of building a unified foundation model that can be adapted to both vision-only and text-only tasks. Starting from BERT and ViT, we design…
cs.CV2021★ 2 cited
Contextualized Spatio-Temporal Contrastive Learning with Self-Supervision
Liangzhe Yuan, Rui Qian, Yin Cui +5
Modern self-supervised learning algorithms typically enforce persistency of instance representations across views. While being very effective on learning holistic image and video r…
cs.MM2014★ 12 cited
Building A Large Concept Bank for Representing Events in Video
Yin Cui, Dong Liu, Jiawei Chen +1
Concept-based video representation has proven to be effective in complex event detection. However, existing methods either manually design concepts or directly adopt concept librar…