VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text
arXiv:2104.11178
Abstract
We present a framework for learning multimodal representations from unlabeled data using convolution-free Transformer architectures. Specifically, our Video-Audio-Text Transformer (VATT) takes raw signals as inputs and extracts multimodal representations that are rich enough to benefit a variety of downstream tasks. We train VATT end-to-end from scratch using multimodal contrastive losses and evaluate its performance by the downstream tasks of video action recognition, audio event classification, image classification, and text-to-video retrieval. Furthermore, we study a modality-agnostic, single-backbone Transformer by sharing weights among the three modalities. We show that the convolution-free VATT outperforms state-of-the-art ConvNet-based architectures in the downstream tasks. Especially, VATT's vision Transformer achieves the top-1 accuracy of 82.1% on Kinetics-400, 83.6% on Kinetics-600, 72.7% on Kinetics-700, and 41.1% on Moments in Time, new records while avoiding supervised pre-training. Transferring to image classification leads to 78.7% top-1 accuracy on ImageNet compared to 64.7% by training the same Transformer from scratch, showing the generalizability of our model despite the domain gap between videos and images. VATT's audio Transformer also sets a new record on waveform-based audio event recognition by achieving the mAP of 39.4% on AudioSet without any supervised pre-training. VATT's source code is publicly available.
Published in the 35th Conference on Neural Information Processing Systems (NeurIPS 2021)
References in corpus (28)
- Adam: A Method for Stochastic Optimization
- Efficient Estimation of Word Representations in Vector Space
- mixup: Beyond Empirical Risk Minimization
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Gaussian Error Linear Units (GELUs)
- Language Models are Few-Shot Learners
- Convolutional Sequence to Sequence Learning
- Unsupervised Learning of Visual Features by Contrasting Cluster Assignments
- Unsupervised Learning of Video Representations using LSTMs
- Is Space-Time Attention All You Need for Video Understanding?
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips
- Data-Efficient Image Recognition with Contrastive Predictive Coding
- Deep Predictive Coding Networks for Video Prediction and Unsupervised Learning
- Cooperative Learning of Audio and Video Models from Self-Supervised Synchronization
- A Short Note about Kinetics-600
- Self-Supervised Learning by Cross-Modal Audio-Video Clustering
- Learning Video Representations using Contrastive Bidirectional Transformer
- Stand-Alone Self-Attention in Vision Models
- Towards Automatic Learning of Procedures from Web Instructional Videos
- Sample-level Deep Convolutional Neural Networks for Music Auto-tagging Using Raw Waveforms
- More Is Less: Learning Efficient Video Representations by Big-Little Network and Depthwise Temporal Aggregation
- Weakly Labelled AudioSet Tagging with Attention Neural Networks
- Memory-augmented Dense Predictive Coding for Video Representation Learning
- Self-supervised Video Representation Learning by Pace Prediction
- An Image is Worth 16x16 Words, What is a Video Worth?
- Learn to cycle: Time-consistent feature discovery for action recognition
- AssembleNet: Searching for Multi-Stream Neural Connectivity in Video Architectures
- -Nets: Double Attention Networks
Cited by in corpus (39)
- On the Opportunities and Risks of Foundation Models
- Human Action Recognition from Various Data Modalities: A Review
- VLP: A Survey on Vision-Language Pre-training
- Attention Bottlenecks for Multimodal Fusion
- Multimodal Co-learning: Challenges, Applications with Datasets, Recent Advances and Future Directions
- Perceiver IO: A General Architecture for Structured Inputs & Outputs
- Video Transformers: A Survey
- Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm
- Self-supervised remote sensing feature learning: Learning Paradigms, Challenges, and Future Works
- When Vision Transformers Outperform ResNets without Pre-training or Strong Data Augmentations
- VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling
- MCare: Learning with Missing Modalities in Multimodal Healthcare Data
- COCOA: Cross Modality Contrastive Learning for Sensor Data
- From CNNs to Transformers in Multimodal Human Action Recognition: A Survey
- MERLOT: Multimodal Neural Script Knowledge Models
- SLAM: A Unified Encoder for Speech and Language Modeling via Speech-Text Joint Pre-Training
- Keeping Your Eye on the Ball: Trajectory Attention in Video Transformers
- LLM-FP4: 4-Bit Floating-Point Quantized Transformers
- SwinCross: Cross-modal Swin Transformer for Head-and-Neck Tumor Segmentation in PET/CT Images
- UFO: A UniFied TransfOrmer for Vision-Language Representation Learning
- Contrastive Learning with Cross-Modal Knowledge Mining for Multimodal Human Activity Recognition
- Learning Vision-Guided Quadrupedal Locomotion End-to-End with Cross-Modal Transformers
- Masked Vision-Language Transformer in Fashion
- Fusion of Satellite Images and Weather Data with Transformer Networks for Downy Mildew Disease Detection
- Self-supervised Graphs for Audio Representation Learning with Limited Labeled Data
- Vid2Seq: Large-Scale Pretraining of a Visual Language Model for Dense Video Captioning
- UAMD-Net: A Unified Adaptive Multimodal Neural Network for Dense Depth Completion
- A vector quantized masked autoencoder for audiovisual speech emotion recognition
- Joint Self-Supervised and Supervised Contrastive Learning for Multimodal MRI Data: Towards Predicting Abnormal Neurodevelopment
- Uni-Perceiver: Pre-training Unified Architecture for Generic Perception for Zero-shot and Few-shot Tasks
- Shifted Chunk Transformer for Spatio-Temporal Representational Learning
- Semantically Guided Representation Learning For Action Anticipation
- Self-supervised Video-centralised Transformer for Video Face Clustering
- Driver-Net: Multi-Camera Fusion for Assessing Driver Take-Over Readiness in Automated Vehicles
- PolyViT: Co-training Vision Transformers on Images, Videos and Audio
- Long-Short Temporal Contrastive Learning of Video Transformers
- Survey: Transformer based Video-Language Pre-training
- Multimodal End-to-End Group Emotion Recognition using Cross-Modal Attention
- Overview of Tencent Multi-modal Ads Video Understanding Challenge