Learning Video Representations from Textual Web Supervision
arXiv:2007.14937
Abstract
Videos on the Internet are paired with pieces of text, such as titles and descriptions. This text typically describes the most important content in the video, such as the objects in the scene and the actions being performed. Based on this observation, we propose to use text as a method for learning video representations. To accomplish this, we propose a data collection process and use it to collect 70M video clips shared publicly on the Internet, and we then train a model to pair each video with its associated text. We evaluate the model on several down-stream action recognition tasks, including Kinetics, HMDB-51, and UCF-101. We find that this approach is an effective method of pre-training video representations. Specifically, it outperforms all existing methods for self-supervised and cross-modal video representation learning.
References in corpus (9)
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- The Kinetics Human Action Video Dataset
- YouTube-8M: A Large-Scale Video Classification Benchmark
- A Short Note on the Kinetics-700-2020 Human Action Dataset
- Spatiotemporal Contrastive Video Representation Learning
- Video Representation Learning with Visual Tempo Consistency
- Watching the World Go By: Representation Learning from Unlabeled Videos
- Learning Spatiotemporal Features via Video and Text Pair Discrimination
- Mining YouTube - A dataset for learning fine-grained action concepts from webly supervised video data
Cited by in corpus (5)
- Learning Transferable Visual Models From Natural Language Supervision
- CLIP2Video: Mastering Video-Text Retrieval via Image CLIP
- Learning Spatiotemporal Features via Video and Text Pair Discrimination
- Revisiting 3D ResNets for Video Recognition
- Revamping Cross-Modal Recipe Retrieval with Hierarchical Transformers and Self-supervised Learning