Attention-Based Models for Speech Recognition
arXiv:1506.07503
Abstract
Recurrent sequence generators conditioned on input data through an attention mechanism have recently shown very good performance on a range of tasks in- cluding machine translation, handwriting synthesis and image caption gen- eration. We extend the attention-mechanism with features needed for speech recognition. We show that while an adaptation of the model used for machine translation in reaches a competitive 18.7% phoneme error rate (PER) on the TIMIT phoneme recognition task, it can only be applied to utterances which are roughly as long as the ones it was trained on. We offer a qualitative explanation of this failure and propose a novel and generic method of adding location-awareness to the attention mechanism to alleviate this issue. The new method yields a model that is robust to long inputs and achieves 18% PER in single utterances and 20% in 10-times longer (repeated) utterances. Finally, we propose a change to the at- tention mechanism that prevents it from concentrating too much on single frames, which further reduces PER to 17.6% level.
References in corpus (10)
- Sequence to Sequence Learning with Neural Networks
- Improving neural networks by preventing co-adaptation of feature detectors
- ADADELTA: An Adaptive Learning Rate Method
- Sequence Transduction with Recurrent Neural Networks
- Theano: new features and speed improvements
- Recurrent Models of Visual Attention
- On Using Monolingual Corpora in Neural Machine Translation
- End-to-end Continuous Speech Recognition using Attention-based Recurrent NN: First Results
- Blocks and Fuel: Frameworks for deep learning
- Neural Turing Machines
Cited by in corpus (149)
- Deep Learning for Audio Signal Processing
- Focusing Attention: Towards Accurate Text Recognition in Natural Images
- Deep Learning Scaling is Predictable, Empirically
- Professor Forcing: A New Algorithm for Training Recurrent Networks
- Stand-Alone Self-Attention in Vision Models
- Reward Augmented Maximum Likelihood for Neural Structured Prediction
- DDGK: Learning Graph Representations for Deep Divergence Graph Kernels
- 2D Attentional Irregular Scene Text Recognizer
- Multichannel End-to-end Speech Recognition
- Res3ATN -- Deep 3D Residual Attention Network for Hand Gesture Recognition in Videos
- Privacy-Preserving Adversarial Representation Learning in ASR: Reality or Illusion?
- Speaker Adaptation for Attention-Based End-to-End Speech Recognition
- How to Teach DNNs to Pay Attention to the Visual Modality in Speech Recognition
- Towards Online End-to-end Transformer Automatic Speech Recognition
- MultiSpeech: Multi-Speaker Text to Speech with Transformer
- Towards Transfer Learning for End-to-End Speech Synthesis from Deep Pre-Trained Language Models
- 3D Gated Recurrent Fusion for Semantic Scene Completion
- Cross-Modal Hierarchical Modelling for Fine-Grained Sketch Based Image Retrieval
- Unsupervised pre-training for sequence to sequence speech recognition
- On the Comparison of Popular End-to-End Models for Large Scale Speech Recognition
- Learning Blended, Precise Semantic Program Embeddings
- A Comparison of Label-Synchronous and Frame-Synchronous End-to-End Models for Speech Recognition
- AReLU: Attention-based Rectified Linear Unit
- A Spatially and Temporally Attentive Joint Trajectory Prediction Framework for Modeling Vessel Intent
- Streaming Chunk-Aware Multihead Attention for Online End-to-End Speech Recognition
- ESPnet-TTS: Unified, Reproducible, and Integratable Open Source End-to-End Text-to-Speech Toolkit
- Replay attack detection with complementary high-resolution information using end-to-end DNN for the ASVspoof 2019 Challenge
- Gated Embeddings in End-to-End Speech Recognition for Conversational-Context Fusion
- Tensor Low-Rank Reconstruction for Semantic Segmentation
- Hard-Coded Gaussian Attention for Neural Machine Translation
- Attention based Convolutional Recurrent Neural Network for Environmental Sound Classification
- Spike-Triggered Non-Autoregressive Transformer for End-to-End Speech Recognition
- Attentive Adversarial Learning for Domain-Invariant Training
- Improving EEG based Continuous Speech Recognition
- Generative Pre-Training for Speech with Autoregressive Predictive Coding
- Saliency-driven Word Alignment Interpretation for Neural Machine Translation
- GASL: Guided Attention for Sparsity Learning in Deep Neural Networks
- end-to-end training of a large vocabulary end-to-end speech recognition system
- From Senones to Chenones: Tied Context-Dependent Graphemes for Hybrid Speech Recognition
- Decoupled Attention Network for Text Recognition
- FastS2S-VC: Streaming Non-Autoregressive Sequence-to-Sequence Voice Conversion
- End-to-End Speech Recognition: A review for the French Language
- Large-Scale Self- and Semi-Supervised Learning for Speech Translation
- Attention and Localization based on a Deep Convolutional Recurrent Model for Weakly Supervised Audio Tagging
- Sparse Sequence-to-Sequence Models
- Effective parameter estimation methods for an ExcitNet model in generative text-to-speech systems
- Weight Pruning via Adaptive Sparsity Loss
- A Novel Fusion of Attention and Sequence to Sequence Autoencoders to Predict Sleepiness From Speech
- Attention-based Transducer for Online Speech Recognition
- ClovaCall: Korean Goal-Oriented Dialog Speech Corpus for Automatic Speech Recognition of Contact Centers
- End-to-End ASR for Code-switched Hindi-English Speech
- Mask CTC: Non-Autoregressive End-to-End ASR with CTC and Mask Predict
- Independent language modeling architecture for end-to-end ASR
- Improving RNN Transducer Modeling for End-to-End Speech Recognition
- End-to-End Spoken Language Translation
- Multilingual End-to-End Speech Translation
- Developing RNN-T Models Surpassing High-Performance Hybrid Models with Customization Capability
- Characterizing Deep Learning Training Workloads on Alibaba-PAI
- Synchronous Transformers for End-to-End Speech Recognition
- Joint Speaker Counting, Speech Recognition, and Speaker Identification for Overlapped Speech of Any Number of Speakers
- Audio2Face: Generating Speech/Face Animation from Single Audio with Attention-Based Bidirectional LSTM Networks
- Distilling the Knowledge of BERT for Sequence-to-Sequence ASR
- Learn Spelling from Teachers: Transferring Knowledge from Language Models to Sequence-to-Sequence Speech Recognition
- Unsupervised Speaker Adaptation using Attention-based Speaker Memory for End-to-End ASR
- Candidate Fusion: Integrating Language Modelling into a Sequence-to-Sequence Handwritten Word Recognition Architecture
- A Transformer with Interleaved Self-attention and Convolution for Hybrid Acoustic Models
- Attention Forcing for Sequence-to-sequence Model Training
- Sequence to Multi-Sequence Learning via Conditional Chain Mapping for Mixture Signals
- Capsule-Transformer for Neural Machine Translation
- Exploring Machine Speech Chain for Domain Adaptation and Few-Shot Speaker Adaptation
- Multi-view Frequency LSTM: An Efficient Frontend for Automatic Speech Recognition
- Effective Sentence Scoring Method using Bidirectional Language Model for Speech Recognition
- End-to-end Anchored Speech Recognition
- An Effective End-to-End Modeling Approach for Mispronunciation Detection
- Guiding CTC Posterior Spike Timings for Improved Posterior Fusion and Knowledge Distillation
- a novel cross-lingual voice cloning approach with a few text-free samples
- Speech-to-speech Translation between Untranscribed Unknown Languages
- Large-Scale Pre-Training of End-to-End Multi-Talker ASR for Meeting Transcription with Single Distant Microphone
- BUT Opensat 2019 Speech Recognition System
- Investigation of End-To-End Speaker-Attributed ASR for Continuous Multi-Talker Recordings
- Pretraining Techniques for Sequence-to-Sequence Voice Conversion
- Transformer-based Online CTC/attention End-to-End Speech Recognition Architecture
- Towards Fast and Accurate Streaming End-to-End ASR
- A Pyramid Recurrent Network for Predicting Crowdsourced Speech-Quality Ratings of Real-World Signals
- AccentDB: A Database of Non-Native English Accents to Assist Neural Speech Recognition
- Learning Frame Level Attention for Environmental Sound Classification
- Adversarial Joint Training with Self-Attention Mechanism for Robust End-to-End Speech Recognition
- Self-and-Mixed Attention Decoder with Deep Acoustic Structure for Transformer-based LVCSR
- An Overview on Data Representation Learning: From Traditional Feature Learning to Recent Deep Learning
- Text-Independent Speaker Verification with Dual Attention Network
- Forward-Backward Decoding for Regularizing End-to-End TTS
- Improving EEG based continuous speech recognition using GAN
- The AS-NU System for the M2VoC Challenge
- Gradual Machine Learning for Aspect-level Sentiment Analysis
- Enhanced 3D Human Pose Estimation from Videos by using Attention-Based Neural Network with Dilated Convolutions
- A Novel Topology for End-to-end Temporal Classification and Segmentation with Recurrent Neural Network
- End-to-End Code-Switching ASR for Low-Resourced Language Pairs
- Unidirectional Memory-Self-Attention Transducer for Online Speech Recognition
- Attention Forcing for Machine Translation
- KoSpeech: Open-Source Toolkit for End-to-End Korean Speech Recognition
- Query-based Interactive Recommendation by Meta-Path and Adapted Attention-GRU
- An Optimized and Energy-Efficient Parallel Implementation of Non-Iteratively Trained Recurrent Neural Networks
- Transformer ASR with Contextual Block Processing
- Attention-based ASR with Lightweight and Dynamic Convolutions
- Domain Adaptation via Teacher-Student Learning for End-to-End Speech Recognition
- Integrating Source-channel and Attention-based Sequence-to-sequence Models for Speech Recognition
- ALCNN: Attention-based Model for Fine-grained Demand Inference of Dock-less Shared Bike in New Cities
- Translate Reverberated Speech to Anechoic Ones: Speech Dereverberation with BERT
- Scaling Up Multiagent Reinforcement Learning for Robotic Systems: Learn an Adaptive Sparse Communication Graph
- Understanding effect of speech perception in EEG based speech recognition systems
- Learning Shared Encoding Representation for End-to-End Speech Recognition Models
- Minimum Latency Training Strategies for Streaming Sequence-to-Sequence ASR
- Homophone-based Label Smoothing in End-to-End Automatic Speech Recognition
- An End-to-End Mispronunciation Detection System for L2 English Speech Leveraging Novel Anti-Phone Modeling
- Exploring Pre-training with Alignments for RNN Transducer based End-to-End Speech Recognition
- Neural Code Summarization
- L-Vector: Neural Label Embedding for Domain Adaptation
- Fast and Accurate Deep Bidirectional Language Representations for Unsupervised Learning
- End-to-End Speech Recognition with High-Frame-Rate Features Extraction
- Improved Speech Representations with Multi-Target Autoregressive Predictive Coding
- Head-synchronous Decoding for Transformer-based Streaming ASR
- Sequence-to-Set Semantic Tagging: End-to-End Multi-label Prediction using Neural Attention for Complex Query Reformulation and Automated Text Categorization
- FSR: Accelerating the Inference Process of Transducer-Based Models by Applying Fast-Skip Regularization
- Small energy masking for improved neural network training for end-to-end speech recognition
- Improving LPCNet-based Text-to-Speech with Linear Prediction-structured Mixture Density Network
- End-to-end Adaptation with Backpropagation through WFST for On-device Speech Recognition System
- Attention based on-device streaming speech recognition with large speech corpus
- LSTM Acoustic Models Learn to Align and Pronounce with Graphemes
- Streaming Multi-talker Speech Recognition with Joint Speaker Identification
- Speaker-aware speech-transformer
- Mutually-Constrained Monotonic Multihead Attention for Online ASR
- Emotional speech synthesis with rich and granularized control
- Two-Headed Monster And Crossed Co-Attention Networks
- A practical two-stage training strategy for multi-stream end-to-end speech recognition
- Multiple-hypothesis CTC-based semi-supervised adaptation of end-to-end speech recognition
- End-to-end Music-mixed Speech Recognition
- Layer-stacked Attention for Heterogeneous Network Embedding
- Far-Field Automatic Speech Recognition
- ProbaNet: Proposal-balanced Network for Object Detection
- CHEER: Rich Model Helps Poor Model via Knowledge Infusion
- Performance Monitoring for End-to-End Speech Recognition
- Improving Accent Conversion with Reference Encoder and End-To-End Text-To-Speech
- Pedestrian Tracking with Gated Recurrent Units and Attention Mechanisms
- Attention-Passing Models for Robust and Data-Efficient End-to-End Speech Translation
- Speaker Adaptation for End-to-End CTC Models
- Improve SGD Training via Aligning Mini-batches
- Exploration of Audio Quality Assessment and Anomaly Localisation Using Attention Models
- Acoustic-to-Word Models with Conversational Context Information
- Character-Aware Attention-Based End-to-End Speech Recognition