Wav2Letter: an End-to-End ConvNet-based Speech Recognition System
arXiv:1609.03193
Abstract
This paper presents a simple end-to-end model for speech recognition, combining a convolutional network based acoustic model and a graph decoding. It is trained to output letters, with transcribed speech, without the need for force alignment of phonemes. We introduce an automatic segmentation criterion for training from sequence annotation without alignment that is on par with CTC while being simpler. We show competitive results in word error rate on the Librispeech corpus with MFCC features, and promising results from raw waveform.
8 pages, 4 figures (7 plots/schemas), 2 tables (4 tabulars)
Cited by in corpus (37)
- MLS: A Large-Scale Multilingual Dataset for Speech Research
- vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations
- Self-Attention Networks for Connectionist Temporal Classification in Speech Recognition
- Towards Goal-Oriented Semantic Signal Processing: Applications and Future Challenges
- Exploring Neural Transducers for End-to-End Speech Recognition
- Dense Prediction on Sequences with Time-Dilated Convolutions for Speech Recognition
- Open-Domain Conversational Agents: Current Progress, Open Problems, and Future Directions
- Why does CTC result in peaky behavior?
- A Fully Differentiable Beam Search Decoder
- Espresso: A Fast End-to-end Neural Speech Recognition Toolkit
- Learning Robust and Multilingual Speech Representations
- Improved Regularization Techniques for End-to-End Speech Recognition
- Reducing Bias in Production Speech Models
- CAT: CRF-based ASR Toolkit
- Attention-based Wav2Text with Feature Transfer Learning
- Differentiable Weighted Finite-State Transducers
- INT8 Winograd Acceleration for Conv1D Equipped ASR Models Deployed on Mobile Devices
- Kaleidoscope: An Efficient, Learnable Representation For All Structured Linear Maps
- Improving End-to-End Speech Recognition with Policy Learning
- Intelligent Video Editing: Incorporating Modern Talking Face Generation Algorithms in a Video Editor
- End-to-End ASR for Code-switched Hindi-English Speech
- Exploring spectro-temporal features in end-to-end convolutional neural networks
- The 2020 ESPnet update: new features, broadened applications, performance improvements, and future plans
- Beyond clipping: Equalization-based Psychoacoustic Attacks against ASRs
- Effects of Number of Filters of Convolutional Layers on Speech Recognition Model Accuracy
- Utilizing Domain Knowledge in End-to-End Audio Processing
- Weakly Supervised Construction of ASR Systems with Massive Video Data
- MASRI-HEADSET: A Maltese Corpus for Speech Recognition
- An Improved Model for Voicing Silent Speech
- Multi-QuartzNet: Multi-Resolution Convolution for Speech Recognition with Multi-Layer Feature Fusion
- CAT: A CTC-CRF based ASR Toolkit Bridging the Hybrid and the End-to-end Approaches towards Data Efficiency and Low Latency
- End-to-End Language Identification using Multi-Head Self-Attention and 1D Convolutional Neural Networks
- CarneliNet: Neural Mixture Model for Automatic Speech Recognition
- Multi-Modal Transformers Utterance-Level Code-Switching Detection
- Exploiting Nontrivial Connectivity for Automatic Speech Recognition
- Densifying Assumed-sparse Tensors: Improving Memory Efficiency and MPI Collective Performance during Tensor Accumulation for Parallelized Training of Neural Machine Translation Models
- Gradient-Adjusted Neuron Activation Profiles for Comprehensive Introspection of Convolutional Speech Recognition Models