Listen, Attend and Spell
arXiv:1508.01211
Abstract
We present Listen, Attend and Spell (LAS), a neural network that learns to transcribe speech utterances to characters. Unlike traditional DNN-HMM models, this model learns all the components of a speech recognizer jointly. Our system has two components: a listener and a speller. The listener is a pyramidal recurrent network encoder that accepts filter bank spectra as inputs. The speller is an attention-based recurrent network decoder that emits characters as outputs. The network produces character sequences without making any independence assumptions between the characters. This is the key improvement of LAS over previous end-to-end CTC models. On a subset of the Google voice search task, LAS achieves a word error rate (WER) of 14.1% without a dictionary or a language model, and 10.3% with language model rescoring over the top 32 beams. By comparison, the state-of-the-art CLDNN-HMM model achieves a WER of 8.0%.
References in corpus (4)
Cited by in corpus (92)
- Deep Speech 2: End-to-End Speech Recognition in English and Mandarin
- Deep Learning for Audio Signal Processing
- Deep Audio-Visual Speech Recognition
- Semi-supervised Sequence Learning
- Feed-Forward Networks with Attention Can Solve Some Long-Term Memory Problems
- Device Placement Optimization with Reinforcement Learning
- Understanding LSTM -- a tutorial into Long Short-Term Memory Recurrent Neural Networks
- Neural Language Correction with Character-Based Attention
- Generating News Headlines with Recurrent Neural Networks
- Survey on the attention based RNN model and its applications in computer vision
- Attention with Intention for a Neural Network Conversation Model
- Reward Augmented Maximum Likelihood for Neural Structured Prediction
- OpenNMT: Neural Machine Translation Toolkit
- Fully Convolutional Speech Recognition
- Sequence-to-sequence neural network models for transliteration
- A Survey on State-of-the-art Deep Learning Applications and Challenges
- Unified Streaming and Non-streaming Two-pass End-to-end Model for Speech Recognition
- Citrinet: Closing the Gap between Non-Autoregressive and Autoregressive End-to-End Models for Automatic Speech Recognition
- Segmental Recurrent Neural Networks for End-to-end Speech Recognition
- Mixed-Precision Training for NLP and Speech Recognition with OpenSeq2Seq
- A Neural Transducer
- Recurrent neural networks and transfer learning for elasto-plasticity in woven composites
- Machine Learning on Sequential Data Using a Recurrent Weighted Average
- Uncertainty Estimation in Autoregressive Structured Prediction
- Unifying Cardiovascular Modelling with Deep Reinforcement Learning for Uncertainty Aware Control of Sepsis Treatment
- Tied & Reduced RNN-T Decoder
- U2++: Unified Two-pass Bidirectional End-to-end Model for Speech Recognition
- Cross-Language Transfer Learning, Continuous Learning, and Domain Adaptation for End-to-End Automatic Speech Recognition
- WeNet: Production oriented Streaming and Non-streaming End-to-End Speech Recognition Toolkit
- Scaling End-to-End Models for Large-Scale Multilingual ASR
- Emphasizing Unseen Words: New Vocabulary Acquisition for End-to-End Speech Recognition
- Multitask-Based Joint Learning Approach To Robust ASR For Radio Communication Speech
- Sparse Attentive Memory Network for Click-through Rate Prediction with Long Sequences
- A Simple Baseline for Domain Adaptation in End to End ASR Systems Using Synthetic Data
- Foreign English Accent Adjustment by Learning Phonetic Patterns
- WNARS: WFST based Non-autoregressive Streaming End-to-End Speech Recognition
- Streaming End-to-End Bilingual ASR Systems with Joint Language Identification
- Attention based end to end Speech Recognition for Voice Search in Hindi and English
- Federated Transfer Learning with Dynamic Gradient Aggregation
- Class LM and word mapping for contextual biasing in End-to-End ASR
- FastEmit: Low-latency Streaming ASR with Sequence-level Emission Regularization
- Attention-based Transducer for Online Speech Recognition
- Simplified End-to-End MMI Training and Voting for ASR
- Hierarchical Sequence to Sequence Voice Conversion with Limited Data
- Contextualizing ASR Lattice Rescoring with Hybrid Pointer Network Language Model
- An Attention-Based Approach for Single Image Super Resolution
- Online Sequence Training of Recurrent Neural Networks with Connectionist Temporal Classification
- Smart at what cost? Characterising Mobile Deep Neural Networks in the wild
- Cross-Attention End-to-End ASR for Two-Party Conversations
- Exploring Machine Speech Chain for Domain Adaptation and Few-Shot Speaker Adaptation
- A Hybrid Vision Transformer Approach for Mathematical Expression Recognition
- Speech2Slot: An End-to-End Knowledge-based Slot Filling from Speech
- A Streaming End-to-End Framework For Spoken Language Understanding
- Convolutional Speech Recognition with Pitch and Voice Quality Features
- Large-Scale Multilingual Speech Recognition with a Streaming End-to-End Model
- Teaching Machines to Code: Neural Markup Generation with Visual Attention
- Neural Composition: Learning to Generate from Multiple Models
- Multi-Speaker ASR Combining Non-Autoregressive Conformer CTC and Conditional Speaker Chain
- Sub-word Level Lip Reading With Visual Attention
- Towards Fast and Accurate Streaming End-to-End ASR
- A Streaming On-Device End-to-End Model Surpassing Server-Side Conventional Model Quality and Latency
- An online sequence-to-sequence model for noisy speech recognition
- Improved Speech Separation with Time-and-Frequency Cross-domain Joint Embedding and Clustering
- Unsupervised Learning of Audio Perception for Robotics Applications: Learning to Project Data to T-SNE/UMAP space
- Back from the future: bidirectional CTC decoding using future information in speech recognition
- Sequence-to-Sequence Modeling for Action Identification at High Temporal Resolution
- A Better and Faster End-to-End Model for Streaming ASR
- Adversarial Joint Training with Self-Attention Mechanism for Robust End-to-End Speech Recognition
- Worse WER, but Better BLEU? Leveraging Word Embedding as Intermediate in Multitask End-to-End Speech Translation
- Chained Predictions Using Convolutional Neural Networks
- Howl: A Deployed, Open-Source Wake Word Detection System
- KoSpeech: Open-Source Toolkit for End-to-End Korean Speech Recognition
- Parallel Rescoring with Transformer for Streaming On-Device Speech Recognition
- Investigation of Training Label Error Impact on RNN-T
- Robot Sound Interpretation: Combining Sight and Sound in Learning-Based Control
- The Performance Evaluation of Attention-Based Neural ASR under Mixed Speech Input
- Exploring Targeted Universal Adversarial Perturbations to End-to-end ASR Models
- Speech Corpus of Ainu Folklore and End-to-end Speech Recognition for Ainu Language
- Multi-mode Transformer Transducer with Stochastic Future Context
- Efficient Transformer for Direct Speech Translation
- Sequence-level self-learning with multiple hypotheses
- Integrating Categorical Features in End-to-End ASR
- Acoustic-to-Word Models with Conversational Context Information
- Speaker Adaptation for End-to-End CTC Models
- Attention based on-device streaming speech recognition with large speech corpus
- Multitask Learning and Joint Optimization for Transformer-RNN-Transducer Speech Recognition
- Quantum Statistics-Inspired Neural Attention
- Master Thesis: Neural Sign Language Translation by Learning Tokenization
- Semantic Data Augmentation for End-to-End Mandarin Speech Recognition
- Loss Prediction: End-to-End Active Learning Approach For Speech Recognition
- Using Deep Learning Sequence Models to Identify SARS-CoV-2 Divergence
- Compositional Sentence Representation from Character within Large Context Text