Efficient Training of Audio Transformers with Patchout
arXiv:2110.05069 · doi:10.21437/Interspeech.2022-227
Abstract
The great success of transformer-based models in natural language processing (NLP) has led to various attempts at adapting these architectures to other domains such as vision and audio. Recent work has shown that transformers can outperform Convolutional Neural Networks (CNNs) on vision and audio tasks. However, one of the main shortcomings of transformer models, compared to the well-established CNNs, is the computational complexity. In transformers, the compute and memory complexity is known to grow quadratically with the input length. Therefore, there has been extensive work on optimizing transformers, but often at the cost of degrading predictive performance. In this work, we propose a novel method to optimize and regularize transformers on audio spectrograms. Our proposed models achieve a new state-of-the-art performance on Audioset and can be trained on a single consumer-grade GPU. Furthermore, we propose a transformer model that outperforms CNNs in terms of both performance and training speed. Source code: https://github.com/kkoutini/PaSST
Submitted to Interspeech 2022. Source code: https://github.com/kkoutini/PaSST
References in corpus (7)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Deformable DETR: Deformable Transformers for End-to-End Object Detection
- Reformer: The Efficient Transformer
- Intriguing Properties of Vision Transformers
- PSLA: Improving Audio Tagging with Pretraining, Sampling, Labeling, and Aggregation
- Perceiver: General Perception with Iterative Attention
- Receptive Field Regularization Techniques for Audio Classification and Tagging with Deep Convolutional Neural Networks
Cited by in corpus (13)
- ASiT: Local-Global Audio Spectrogram vIsion Transformer for Event Classification
- Sounding Video Generator: A Unified Framework for Text-guided Sounding Video Generation
- Benchmarking Representations for Speech, Music, and Acoustic Events
- LEAN: Light and Efficient Audio Classification Network
- Spectrogram features for audio and speech analysis
- Recent Advances and Challenges in Deep Audio-Visual Correlation Learning
- Masked Latent Prediction and Classification for Self-Supervised Audio Representation Learning
- Data-Efficient Low-Complexity Acoustic Scene Classification via Distilling and Progressive Pruning
- SMITIN: Self-Monitored Inference-Time INtervention for Generative Music Transformers
- UniKW-AT: Unified Keyword Spotting and Audio Tagging
- From Physics to Representation: Audio Learning with Synthetic Pre-training via Procedural Generation
- Hierarchical Label Propagation: A Model-Size-Dependent Performance Booster for AudioSet Tagging
- Masked Autoencoders with Multi-Window Local-Global Attention Are Better Audio Learners