Keyword Transformer: A Self-Attention Model for Keyword Spotting
arXiv:2104.00769 · doi:10.21437/Interspeech.2021-1286
Abstract
The Transformer architecture has been successful across many domains, including natural language processing, computer vision and speech recognition. In keyword spotting, self-attention has primarily been used on top of convolutional or recurrent encoders. We investigate a range of ways to adapt the Transformer architecture to keyword spotting and introduce the Keyword Transformer (KWT), a fully self-attentional architecture that exceeds state-of-the-art performance across multiple tasks without any pre-training or additional data. Surprisingly, this simple architecture outperforms more complex models that mix convolutional, recurrent and attentive layers. KWT can be used as a drop-in replacement for these models, setting two new benchmark records on the Google Speech Commands dataset with 98.6% and 97.7% accuracy on the 12 and 35-command tasks respectively.
Proceedings of INTERSPEECH
References in corpus (10)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Distilling the Knowledge in a Neural Network
- SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning
- Conformer: Convolution-augmented Transformer for Speech Recognition
- TERA: Self-Supervised Learning of Transformer Encoder Representation for Speech
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices
- Deep Coevolutionary Network: Embedding User and Item Features for Recommendation
- Learning Efficient Representations for Keyword Spotting with Triplet Loss
- Ternary Hybrid Neural-Tree Networks for Highly Constrained IoT Applications
- baller2vec: A Multi-Entity Transformer For Multi-Agent Spatiotemporal Modeling
Cited by in corpus (7)
- Action Transformer: A Self-Attention Model for Short-Time Pose-Based Human Action Recognition
- ConvMixer: Feature Interactive Convolution with Curriculum Learning for Small Footprint and Noisy Far-field Keyword Spotting
- Marvin: an Innovative Omni-Directional Robotic Assistant for Domestic Environments
- Visual Keyword Spotting with Attention
- Attention-Free Keyword Spotting
- Audiomer: A Convolutional Transformer For Keyword Spotting
- Seed Words Based Data Selection for Language Model Adaptation