CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval
arXiv:2104.08860
Abstract
Video-text retrieval plays an essential role in multi-modal research and has been widely used in many real-world web applications. The CLIP (Contrastive Language-Image Pre-training), an image-language pre-training model, has demonstrated the power of visual concepts learning from web collected image-text datasets. In this paper, we propose a CLIP4Clip model to transfer the knowledge of the CLIP model to video-language retrieval in an end-to-end manner. Several questions are investigated via empirical studies: 1) Whether image feature is enough for video-text retrieval? 2) How a post-pretraining on a large-scale video-text dataset based on the CLIP affect the performance? 3) What is the practical mechanism to model temporal dependency between video frames? And 4) The Hyper-parameters sensitivity of the model on video-text retrieval task. Extensive experimental results present that the CLIP4Clip model transferred from the CLIP can achieve SOTA results on various video-text retrieval datasets, including MSR-VTT, MSVC, LSMDC, ActivityNet, and DiDeMo. We release our code at https://github.com/ArrowLuo/CLIP4Clip.
References in corpus (7)
- Learning Transferable Visual Models From Natural Language Supervision
- Microsoft COCO Captions: Data Collection and Evaluation Server
- Is Space-Time Attention All You Need for Video Understanding?
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- MDMMT: Multidomain Multimodal Transformer for Video Retrieval
- Learning Language-Visual Embedding for Movie Understanding with Natural-Language
- HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training
Cited by in corpus (22)
- CLIP2Video: Mastering Video-Text Retrieval via Image CLIP
- CenterCLIP: Token Clustering for Efficient Text-Video Retrieval
- VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling
- Audio Retrieval with Natural Language Queries: A Benchmark Study
- Improving Video-Text Retrieval by Multi-Stream Corpus Alignment and Dual Softmax Loss
- VALUE: A Multi-Task Benchmark for Video-and-Language Understanding Evaluation
- CLOOB: Modern Hopfield Networks with InfoLOOB Outperform CLIP
- Cross-Lingual Cross-Modal Retrieval with Noise-Robust Learning
- Cross-Modal Adapter for Vision-Language Retrieval
- Learn to Understand Negation in Video Retrieval
- ScaleVLAD: Improving Multimodal Sentiment Analysis via Multi-Scale Fusion of Locally Descriptors
- Induce, Edit, Retrieve: Language Grounded Multimodal Schema for Instructional Video Retrieval
- Uni-Perceiver: Pre-training Unified Architecture for Generic Perception for Zero-shot and Few-shot Tasks
- A CLIP-Enhanced Method for Video-Language Understanding
- IDEA: Increasing Text Diversity via Online Multi-Label Recognition for Vision-Language Pre-training
- VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
- STOA-VLP: Spatial-Temporal Modeling of Object and Action for Video-Language Pre-training
- SNP-S3: Shared Network Pre-training and Significant Semantic Strengthening for Various Video-Text Tasks
- ViSeRet: A simple yet effective approach to moment retrieval via fine-grained video segmentation
- EVOQUER: Enhancing Temporal Grounding with Video-Pivoted BackQuery Generation
- CLOP: Video-and-Language Pre-Training with Knowledge Regularizations
- CLIP4Caption: CLIP for Video Caption