Long Short-Term Memory Based Recurrent Neural Network Architectures for Large Vocabulary Speech Recognition
arXiv:1402.1128
Abstract
Long Short-Term Memory (LSTM) is a recurrent neural network (RNN) architecture that has been designed to address the vanishing and exploding gradient problems of conventional RNNs. Unlike feedforward neural networks, RNNs have cyclic connections making them powerful for modeling sequences. They have been successfully used for sequence labeling and sequence prediction tasks, such as handwriting recognition, language modeling, phonetic labeling of acoustic frames. However, in contrast to the deep neural networks, the use of RNNs in speech recognition has been limited to phone recognition in small scale tasks. In this paper, we present novel LSTM based RNN architectures which make more effective use of model parameters to train acoustic models for large vocabulary speech recognition. We train and compare LSTM, RNN and DNN models at various numbers of parameters and configurations. We show that LSTM models converge quickly and give state of the art speech recognition performance for relatively small sized models.
Cited by in corpus (87)
- DeepSleepNet: a Model for Automatic Sleep Stage Scoring based on Raw Single-Channel EEG
- Recent Trends in Deep Learning Based Natural Language Processing
- Recent Advances in Recurrent Neural Networks
- Machine Learning in Aerodynamic Shape Optimization
- Spatio-temporal video autoencoder with differentiable memory
- Contextual LSTM (CLSTM) models for Large scale NLP tasks
- A Comprehensive Study of Deep Bidirectional LSTM RNNs for Acoustic Modeling in Speech Recognition
- Fast Gradient Attack on Network Embedding
- Temporal Localization of Fine-Grained Actions in Videos by Domain Transfer from Web Images
- Going in circles is the way forward: the role of recurrence in visual inference
- Few-Shot Deep Adversarial Learning for Video-based Person Re-identification
- Using Recurrent Neural Networks to Optimize Dynamical Decoupling for Quantum Memory
- Feedforward Sequential Memory Networks: A New Structure to Learn Long-term Dependency
- Sequence-to-Sequence Neural Net Models for Grapheme-to-Phoneme Conversion
- Stacking-Based Deep Neural Network: Deep Analytic Network for Pattern Classification
- StorSeismic: A new paradigm in deep learning for seismic processing
- Forecasting Economics and Financial Time Series: ARIMA vs. LSTM
- Development and Evaluation of Recurrent Neural Network based Models for Hourly Traffic Volume and AADT Prediction
- On the State of the Art of Evaluation in Neural Language Models
- Generative Models for Effective ML on Private, Decentralized Datasets
- LightRNN: Memory and Computation-Efficient Recurrent Neural Networks
- Discriminative Acoustic Word Embeddings: Recurrent Neural Network-Based Approaches
- Learning Quantum Hamiltonians from Single-qubit Measurements
- Fine-Grained Trajectory-based Travel Time Estimation for Multi-city Scenarios Based on Deep Meta-Learning
- FortuneTeller: Predicting Microarchitectural Attacks via Unsupervised Deep Learning
- Deep Learning Based Simulators for the Phosphorus Removal Process Control in Wastewater Treatment via Deep Reinforcement Learning Algorithms
- Leveraging native language information for improved accented speech recognition
- TRANS-BLSTM: Transformer with Bidirectional LSTM for Language Understanding
- Can Adversarial Network Attack be Defended?
- A deep learning approach for lower back-pain risk prediction during manual lifting
- Scalable Learning With a Structural Recurrent Neural Network for Short-Term Traffic Prediction
- Mitigating Edge Machine Learning Inference Bottlenecks: An Empirical Study on Accelerating Google Edge Models
- Sequence Model Design for Code Completion in the Modern IDE
- Molecular Identification from AFM images using the IUPAC Nomenclature and Attribute Multimodal Recurrent Neural Networks
- Understanding Unintended Memorization in Federated Learning
- Stock Movement Prediction with Multimodal Stable Fusion via Gated Cross-Attention Mechanism
- Training Production Language Models without Memorizing User Data
- Long Short-Term Memory Networks for CSI300 Volatility Prediction with Baidu Search Volume
- Constructing Long Short-Term Memory based Deep Recurrent Neural Networks for Large Vocabulary Speech Recognition
- Top-down Tree Long Short-Term Memory Networks
- A limited-size ensemble of homogeneous CNN/LSTMs for high-performance word classification
- Feature-weighted Stacking for Nonseasonal Time Series Forecasts: A Case Study of the COVID-19 Epidemic Curves
- Sparse Attentive Backtracking: Long-Range Credit Assignment in Recurrent Networks
- Capturing Popularity Trends: A Simplistic Non-Personalized Approach for Enhanced Item Recommendation
- Proactive Resource Management for LTE in Unlicensed Spectrum: A Deep Learning Perspective
- Beyond Human-Level Accuracy: Computational Challenges in Deep Learning
- Recurrent Models for Auditory Attention in Multi-Microphone Distance Speech Recognition
- Blending LSTMs into CNNs
- Towards Playlist Generation Algorithms Using RNNs Trained on Within-Track Transitions
- Trusted Artificial Intelligence: Towards Certification of Machine Learning Applications
- Two-pass Decoding and Cross-adaptation Based System Combination of End-to-end Conformer and Hybrid TDNN ASR Systems
- Adversarial Training for Multilingual Acoustic Modeling
- Essence Knowledge Distillation for Speech Recognition
- Deep Canonically Correlated LSTMs
- On the efficient representation and execution of deep acoustic models
- Long Short-Term Memory based Convolutional Recurrent Neural Networks for Large Vocabulary Speech Recognition
- MomentumRNN: Integrating Momentum into Recurrent Neural Networks
- A Review of Intelligent Practices for Irrigation Prediction
- Empirical Evaluation of Speaker Adaptation on DNN based Acoustic Model
- Gated Recurrent Unit Based Acoustic Modeling with Future Context
- Log Message Anomaly Detection and Classification Using Auto-B/LSTM and Auto-GRU
- LSTM-TDNN with convolutional front-end for Dialect Identification in the 2019 Multi-Genre Broadcast Challenge
- Simplified Gating in Long Short-term Memory (LSTM) Recurrent Neural Networks
- Distilling Knowledge Using Parallel Data for Far-field Speech Recognition
- FineText: Text Classification via Attention-based Language Model Fine-tuning
- Multi-task Recurrent Model for True Multilingual Speech Recognition
- Deep learning for brake squeal: vibration detection, characterization and prediction
- Vision-based Navigation of Autonomous Vehicle in Roadway Environments with Unexpected Hazards
- Efficient Contextual Representation Learning Without Softmax Layer
- Manual Post-editing of Automatically Transcribed Speeches from the Icelandic Parliament - Althingi
- Development of Automatic Speech Recognition for Kazakh Language using Transfer Learning
- Regime Learning for Differentiable Particle Filters
- Recycle deep features for better object detection
- Deep Feed-forward Sequential Memory Networks for Speech Synthesis
- Optimizing Block-Sparse Matrix Multiplications on CUDA with TVM
- Improving Gated Recurrent Unit Based Acoustic Modeling with Batch Normalization and Enlarged Context
- Advancing Multi-Accented LSTM-CTC Speech Recognition using a Domain Specific Student-Teacher Learning Paradigm
- Language Recognition using Time Delay Deep Neural Network
- Homophone-based Label Smoothing in End-to-End Automatic Speech Recognition
- Empirical Evaluation of Parallel Training Algorithms on Acoustic Modeling
- Empirical Evaluation of A New Approach to Simplifying Long Short-term Memory (LSTM)
- Efficient Inference via Universal LSH Kernel
- Forex Trading Volatility Prediction using Neural Network Models
- C-DLinkNet: considering multi-level semantic features for human parsing
- Superposition as Data Augmentation using LSTM and HMM in Small Training Sets
- A Morpho-Syntactically Informed LSTM-CRF Model for Named Entity Recognition
- Semi-supervised Learning for Convolutional Neural Networks via Online Graph Construction