Unified Streaming and Non-streaming Two-pass End-to-end Model for Speech Recognition
arXiv:2012.05481
Abstract
In this paper, we present a novel two-pass approach to unify streaming and non-streaming end-to-end (E2E) speech recognition in a single model. Our model adopts the hybrid CTC/attention architecture, in which the conformer layers in the encoder are modified. We propose a dynamic chunk-based attention strategy to allow arbitrary right context length. At inference time, the CTC decoder generates n-best hypotheses in a streaming way. The inference latency could be easily controlled by only changing the chunk size. The CTC hypotheses are then rescored by the attention decoder to get the final result. This efficient rescoring process causes very little sentence-level latency. Our experiments on the open 170-hour AISHELL-1 dataset show that, the proposed method can unify the streaming and non-streaming model simply and efficiently. On the AISHELL-1 test set, our unified model achieves 5.60% relative character error rate (CER) reduction in non-streaming ASR compared to a standard non-streaming transformer. The same model achieves 5.42% CER with 640ms latency in a streaming ASR system.
References in corpus (6)
- Attention-Based Models for Speech Recognition
- Sequence Transduction with Recurrent Neural Networks
- End-to-end Continuous Speech Recognition using Attention-based Recurrent NN: First Results
- Conformer: Convolution-augmented Transformer for Speech Recognition
- Transformer Transducer: One Model Unifying Streaming and Non-streaming Speech Recognition
- Streaming Chunk-Aware Multihead Attention for Online End-to-End Speech Recognition
Cited by in corpus (11)
- Citrinet: Closing the Gap between Non-Autoregressive and Autoregressive End-to-End Models for Automatic Speech Recognition
- U2++: Unified Two-pass Bidirectional End-to-end Model for Speech Recognition
- WeNet: Production oriented Streaming and Non-streaming End-to-End Speech Recognition Toolkit
- ZeroPrompt: Streaming Acoustic Encoders are Zero-Shot Masked LMs
- Emphasizing Unseen Words: New Vocabulary Acquisition for End-to-End Speech Recognition
- WNARS: WFST based Non-autoregressive Streaming End-to-End Speech Recognition
- E2E-based Multi-task Learning Approach to Joint Speech and Accent Recognition
- UCorrect: An Unsupervised Framework for Automatic Speech Recognition Error Correction
- SimulLR: Simultaneous Lip Reading Transducer with Attention-Guided Adaptive Memory
- Decoupling recognition and transcription in Mandarin ASR
- Fast-MD: Fast Multi-Decoder End-to-End Speech Translation with Non-Autoregressive Hidden Intermediates