Who Needs Words? Lexicon-Free Speech Recognition
arXiv:1904.04479 · doi:10.21437/Interspeech.2019-3107
Abstract
Lexicon-free speech recognition naturally deals with the problem of out-of-vocabulary (OOV) words. In this paper, we show that character-based language models (LM) can perform as well as word-based LMs for speech recognition, in word error rates (WER), even without restricting the decoding to a lexicon. We study character-based LMs and show that convolutional LMs can effectively leverage large (character) contexts, which is key for good speech recognition performance downstream. We specifically show that the lexicon-free decoding performance (WER) on utterances with OOV words using character-based LMs is better than lexicon-based decoding, both with character or word-based LMs.
8 pages, 1 figure
References in corpus (9)
- On the difficulty of training Recurrent Neural Networks
- Convolutional Sequence to Sequence Learning
- Improved training of end-to-end attention models for speech recognition
- Wav2Letter: an End-to-End ConvNet-based Speech Recognition System
- wav2letter++: The Fastest Open-source Speech Recognition System
- Fully Convolutional Speech Recognition
- The CAPIO 2017 Conversational Speech Recognition System
- Towards End-to-End Speech Recognition with Deep Convolutional Neural Networks
- Deep Recurrent Neural Networks for Acoustic Modelling
Cited by in corpus (9)
- vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations
- End-to-end ASR: from Supervised to Semi-Supervised Learning with Modern Architectures
- Espresso: A Fast End-to-end Neural Speech Recognition Toolkit
- Emphasizing Unseen Words: New Vocabulary Acquisition for End-to-End Speech Recognition
- HarperValleyBank: A Domain-Specific Spoken Dialog Corpus
- General-Purpose Speech Representation Learning through a Self-Supervised Multi-Granularity Framework
- End-to-end Whispered Speech Recognition with Frequency-weighted Approaches and Pseudo Whisper Pre-training
- PyChain: A Fully Parallelized PyTorch Implementation of LF-MMI for End-to-End ASR
- G2G: TTS-Driven Pronunciation Learning for Graphemic Hybrid ASR