Efficient End-to-End Speech Recognition Using Performers in Conformers
arXiv:2011.04196
Abstract
On-device end-to-end speech recognition poses a high requirement on model efficiency. Most prior works improve the efficiency by reducing model sizes. We propose to reduce the complexity of model architectures in addition to model sizes. More specifically, we reduce the floating-point operations in conformer by replacing the transformer module with a performer. The proposed attention-based efficient end-to-end speech recognition model yields competitive performance on the LibriSpeech corpus with 10 millions of parameters and linear computation complexity. The proposed model also outperforms previous lightweight end-to-end models by about 20% relatively in word error rate.
The current submission has not been reviewed by the coauthor
References in corpus (6)
- Conformer: Convolution-augmented Transformer for Speech Recognition
- ContextNet: Improving Convolutional Neural Networks for Automatic Speech Recognition with Global Context
- Transformer-Transducer: End-to-End Speech Recognition with Self-Attention
- Recent Developments on ESPnet Toolkit Boosted by Conformer
- Transformer Transducer: A Streamable Speech Recognition Model with Transformer Encoders and RNN-T Loss
- Bridging the Gap Between Monaural Speech Enhancement and Recognition with Distortion-Independent Acoustic Modeling