Recent Advances in End-to-End Automatic Speech Recognition
arXiv:2111.01690
Abstract
Recently, the speech community is seeing a significant trend of moving from deep neural network based hybrid modeling to end-to-end (E2E) modeling for automatic speech recognition (ASR). While E2E models achieve the state-of-the-art results in most benchmarks in terms of ASR accuracy, hybrid models are still used in a large proportion of commercial ASR systems at the current time. There are lots of practical factors that affect the production model deployment decision. Traditional hybrid models, being optimized for production for decades, are usually good at these factors. Without providing excellent solutions to all these factors, it is hard for E2E models to be widely commercialized. In this paper, we will overview the recent advances in E2E models, focusing on technologies addressing those challenges from the industry's perspective.
Accepted at APSIPA Transactions on Signal and Information Processing
Cited by in corpus (8)
- Competent but Rigid: Identifying the Gap in Empowering AI to Participate Equally in Group Decision-Making
- A CTC Alignment-based Non-autoregressive Transformer for End-to-end Automatic Speech Recognition
- SAPIEN: Affective Virtual Agents Powered by Large Language Models
- MM-ALT: A Multimodal Automatic Lyric Transcription System
- Text-only domain adaptation for end-to-end ASR using integrated text-to-mel-spectrogram generator
- Fake the Real: Backdoor Attack on Deep Speech Classification via Voice Conversion
- Reducing Geographic Disparities in Automatic Speech Recognition via Elastic Weight Consolidation
- End-to-end Joint Punctuated and Normalized ASR with a Limited Amount of Punctuated Training Data