Attention-Free Keyword Spotting
arXiv:2110.07749
Abstract
Till now, attention-based models have been used with great success in the keyword spotting problem domain. However, in light of recent advances in deep learning, the question arises whether self-attention is truly irreplaceable for recognizing speech keywords. We thus explore the usage of gated MLPs --previously shown to be alternatives to transformers in vision tasks-- for the keyword spotting task. We provide a family of highly efficient MLP-based models for keyword spotting, with less than 0.5 million parameters. We show that our approach achieves competitive performance on Google Speech Commands V2-12 and V2-35 benchmarks with much fewer parameters than self-attention-based methods.
5 pages: Accepted at PML4DC workshop in ICLR 2022
References in corpus (12)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Distilling the Knowledge in a Neural Network
- MLP-Mixer: An all-MLP Architecture for Vision
- Conformer: Convolution-augmented Transformer for Speech Recognition
- TERA: Self-Supervised Learning of Transformer Encoder Representation for Speech
- Keyword Transformer: A Self-Attention Model for Keyword Spotting
- Streaming keyword spotting on mobile devices
- Do You Even Need Attention? A Stack of Feed-Forward Layers Does Surprisingly Well on ImageNet
- Learning Efficient Representations for Keyword Spotting with Triplet Loss
- AST: Audio Spectrogram Transformer
- Pay Attention to MLPs
- MatchboxNet: 1D Time-Channel Separable Convolutional Neural Network Architecture for Speech Commands Recognition