13 citations · 13 across the 1 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2023
Streaming Video Model
Yucheng Zhao, Chong Luo, Chuanxin Tang +3
Video understanding tasks have traditionally been modeled by two separate architectures, specially tailored for two distinct tasks. Sequence-based video tasks, such as action recog…
cs.CV2022★ 4 cited
Learning Visual Representation from Modality-Shared Contrastive Language-Image Pre-training
Haoxuan You, Luowei Zhou, Bin Xiao +5
Large-scale multi-modal contrastive pre-training has demonstrated great utility to learn transferable features for a range of downstream tasks by mapping multiple modalities into a…
cs.CV2021★ 13 cited
RegionCLIP: Region-based Language-Image Pretraining
Yiwu Zhong, Jianwei Yang, Pengchuan Zhang +8
Contrastive language-image pretraining (CLIP) using image-text pairs has achieved impressive results on image classification in both zero-shot and transfer learning settings. Howev…