PolyViT: Co-training Vision Transformers on Images, Videos and Audio
arXiv:2111.12993
Abstract
Can we train a single transformer model capable of processing multiple modalities and datasets, whilst sharing almost all of its learnable parameters? We present PolyViT, a model trained on image, audio and video which answers this question. By co-training different tasks on a single modality, we are able to improve the accuracy of each individual task and achieve state-of-the-art results on 5 standard video- and audio-classification datasets. Co-training PolyViT on multiple modalities and tasks leads to a model that is even more parameter-efficient, and learns representations that generalize across multiple domains. Moreover, we show that co-training is simple and practical to implement, as we do not need to tune hyperparameters for each combination of datasets, but can simply adapt those from standard, single-task training.
References in corpus (12)
- Language Models are Few-Shot Learners
- The Kinetics Human Action Video Dataset
- Is Space-Time Attention All You Need for Video Understanding?
- VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text
- Multi-Task Learning as Multi-Objective Optimization
- One Model To Learn Them All
- A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark
- Perceiver: General Perception with Iterative Attention
- More Is Less: Learning Efficient Video Representations by Big-Little Network and Depthwise Temporal Aggregation
- VGGSound: A Large-scale Audio-Visual Dataset
- SCENIC: A JAX Library for Computer Vision Research and Beyond
- UniT: Multimodal Multitask Learning with a Unified Transformer