A Comparative Study on Transformer vs RNN in Speech Applications
arXiv:1909.06317 · doi:10.1109/ASRU46091.2019.9003750
Abstract
Sequence-to-sequence models have been widely used in end-to-end speech processing, for example, automatic speech recognition (ASR), speech translation (ST), and text-to-speech (TTS). This paper focuses on an emergent sequence-to-sequence model called Transformer, which achieves state-of-the-art performance in neural machine translation and other natural language processing applications. We undertook intensive studies in which we experimentally compared and analyzed Transformer and conventional recurrent neural networks (RNN) in a total of 15 ASR, one multilingual ASR, one ST, and two TTS benchmarks. Our experiments revealed various training tips and significant performance benefits obtained with Transformer for each task including the surprising superiority of Transformer in 13/15 ASR benchmarks in comparison with RNN. We are preparing to release Kaldi-style reproducible recipes using open source and publicly available datasets for all the ASR, ST, and TTS tasks for the community to succeed our exciting outcomes.
Accepted at ASRU 2019
References in corpus (7)
- Sequence to Sequence Learning with Neural Networks
- ADADELTA: An Adaptive Learning Rate Method
- SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition
- FastSpeech: Fast, Robust and Controllable Text to Speech
- Language Modeling with Deep Transformers
- RWTH ASR Systems for LibriSpeech: Hybrid vs Attention -- w/o Data Augmentation
- JSUT corpus: free large-scale Japanese speech corpus for end-to-end speech synthesis
Cited by in corpus (172)
- SpeechBrain: A General-Purpose Speech Toolkit
- Conformer: Convolution-augmented Transformer for Speech Recognition
- Transformer-based Acoustic Modeling for Hybrid Speech Recognition
- GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
- End-to-end ASR: from Supervised to Semi-Supervised Learning with Modern Architectures
- A Practical Survey on Faster and Lighter Transformers
- Improving Transformer-based Speech Recognition Using Unsupervised Pre-training
- Listen and Fill in the Missing Letters: Non-Autoregressive Transformer for Speech Recognition
- CTC-Segmentation of Large Corpora for German End-to-end Speech Recognition
- ContextNet: Improving Convolutional Neural Networks for Automatic Speech Recognition with Global Context
- Online Hybrid CTC/Attention End-to-End Automatic Speech Recognition Architecture
- Citrinet: Closing the Gap between Non-Autoregressive and Autoregressive End-to-End Models for Automatic Speech Recognition
- A New Training Pipeline for an Improved Neural Transducer
- Recent Developments on ESPnet Toolkit Boosted by Conformer
- Transformer-based End-to-End Speech Recognition with Local Dense Synthesizer Attention
- Voice Transformer Network: Sequence-to-Sequence Voice Conversion Using Transformer with Text-to-Speech Pretraining
- The Jazz Transformer on the Front Line: Exploring the Shortcomings of AI-composed Music through Quantitative Measures
- An Overview of Indian Spoken Language Recognition from Machine Learning Perspective
- You Do Not Need More Data: Improving End-To-End Speech Recognition by Text-To-Speech Data Augmentation
- End-to-End Far-Field Speech Recognition with Unified Dereverberation and Beamforming
- AI Choreographer: Music Conditioned 3D Dance Generation with AIST++
- Towards Online End-to-end Transformer Automatic Speech Recognition
- Data Augmentation for End-to-end Code-switching Speech Recognition
- User Retention-oriented Recommendation with Decision Transformer
- Non-autoregressive Transformer-based End-to-end ASR using BERT
- Iterative Pseudo-Labeling for Speech Recognition
- Group Communication with Context Codec for Lightweight Source Separation
- A Text-guided Protein Design Framework
- Semantic Mask for Transformer based End-to-End Speech Recognition
- Improving Generalization of Transformer for Speech Recognition with Parallel Schedule Sampling and Relative Positional Embedding
- Layer-wise Fast Adaptation for End-to-End Multi-Accent Speech Recognition
- Low Latency End-to-End Streaming Speech Recognition with a Scout Network
- The Dark Side of Dataset Scaling: Evaluating Racial Classification in Multimodal Models
- Bayesian Learning for Deep Neural Network Adaptation
- Ensemble of ACCDOA- and EINV2-based Systems with D3Nets and Impulse Response Simulation for Sound Event Localization and Detection
- Towards a Competitive End-to-End Speech Recognition for CHiME-6 Dinner Party Transcription
- A Simplified Fully Quantized Transformer for End-to-end Speech Recognition
- On-the-Fly Aligned Data Augmentation for Sequence-to-Sequence ASR
- On the Comparison of Popular End-to-End Models for Large Scale Speech Recognition
- Mixed Precision Low-bit Quantization of Neural Network Language Models for Speech Recognition
- A Comparison of Label-Synchronous and Frame-Synchronous End-to-End Models for Speech Recognition
- Exploring Transformers for Large-Scale Speech Recognition
- Speech SIMCLR: Combining Contrastive and Reconstruction Objective for Self-supervised Speech Representation Learning
- Multi-head Monotonic Chunkwise Attention For Online Speech Recognition
- Emformer: Efficient Memory Transformer Based Acoustic Model For Low Latency Streaming Speech Recognition
- Streaming Chunk-Aware Multihead Attention for Online End-to-End Speech Recognition
- ESPnet-TTS: Unified, Reproducible, and Integratable Open Source End-to-End Text-to-Speech Toolkit
- Enforcing Encoder-Decoder Modularity in Sequence-to-Sequence Models
- An investigation of phone-based subword units for end-to-end speech recognition
- Multitask-Based Joint Learning Approach To Robust ASR For Radio Communication Speech
- Attention is All You Need in Speech Separation
- Distinguishing a planetary transit from false positives: a Transformer-based classification for planetary transit signals
- Adversarial Attacks and Defenses for Speech Recognition Systems
- A Further Study of Unsupervised Pre-training for Transformer Based Speech Recognition
- SAN-M: Memory Equipped Self-Attention for End-to-End Speech Recognition
- Low-Latency Sequence-to-Sequence Speech Recognition and Translation by Partial Hypothesis Selection
- Developing Real-time Streaming Transformer Transducer for Speech Recognition on Large-scale Dataset
- Exploration of End-to-End ASR for OpenSTT -- Russian Open Speech-to-Text Dataset
- Advanced Long-Content Speech Recognition With Factorized Neural Transducer
- Attention based end to end Speech Recognition for Voice Search in Hindi and English
- Handwritten Mathematical Expression Recognition with Bidirectionally Trained Transformer
- Speech Enhancement using Self-Adaptation and Multi-Head Self-Attention
- Fast End-to-End Speech Recognition via Non-Autoregressive Models and Cross-Modal Knowledge Transferring from BERT
- Intermediate Loss Regularization for CTC-based Speech Recognition
- Leveraging End-to-End ASR for Endangered Language Documentation: An Empirical Study on Yoloxóchitl Mixtec
- BiQGEMM: Matrix Multiplication with Lookup Table For Binary-Coding-based Quantized DNNs
- Two-pass Decoding and Cross-adaptation Based System Combination of End-to-end Conformer and Hybrid TDNN ASR Systems
- Mask CTC: Non-Autoregressive End-to-End ASR with CTC and Mask Predict
- Towards Semi-Supervised Semantics Understanding from Speech
- Improved Mask-CTC for Non-Autoregressive End-to-End ASR
- Speech enhancement with frequency domain auto-regressive modeling
- SpeechNet: A Universal Modularized Model for Speech Processing Tasks
- The 2020 ESPnet update: new features, broadened applications, performance improvements, and future plans
- End-to-End Multi-speaker Speech Recognition with Transformer
- Listen Attentively, and Spell Once: Whole Sentence Generation via a Non-Autoregressive Architecture for Low-Latency Speech Recognition
- Weak-Attention Suppression For Transformer Based Speech Recognition
- A systematic comparison of grapheme-based vs. phoneme-based label units for encoder-decoder-attention models
- Non-autoregressive End-to-end Speech Translation with Parallel Autoregressive Rescoring
- Exoplanet Transit Candidate Identification in TESS Full-Frame Images via a Transformer-Based Algorithm
- SpecAugment on Large Scale Datasets
- Fast Interleaved Bidirectional Sequence Generation
- Advancing CTC-CRF Based End-to-End Speech Recognition with Wordpieces and Conformers
- A Transformer with Interleaved Self-attention and Convolution for Hybrid Acoustic Models
- Momentum Pseudo-Labeling for Semi-Supervised Speech Recognition
- Directional ASR: A New Paradigm for E2E Multi-Speaker Speech Recognition with Source Localization
- Non-Autoregressive Transformer ASR with CTC-Enhanced Decoder Input
- An evaluation of word-level confidence estimation for end-to-end automatic speech recognition
- High-Fidelity Cellular Network Control-Plane Traffic Generation without Domain Knowledge
- Multi-Encoder-Decoder Transformer for Code-Switching Speech Recognition
- On the Usefulness of Self-Attention for Automatic Speech Recognition with Transformers
- Reservoir Network with Structural Plasticity for Human Activity Recognition
- Insertion-Based Modeling for End-to-End Automatic Speech Recognition
- Transformer-based Online CTC/attention End-to-End Speech Recognition Architecture
- Pretraining Techniques for Sequence-to-Sequence Voice Conversion
- End-to-end lyrics Recognition with Voice to Singing Style Transfer
- Don't shoot butterfly with rifles: Multi-channel Continuous Speech Separation with Early Exit Transformer
- Transformer-based Online Speech Recognition with Decoder-end Adaptive Computation Steps
- Wake Word Detection with Streaming Transformers
- Multi-Speaker ASR Combining Non-Autoregressive Conformer CTC and Conditional Speaker Chain
- Multilingual Speech Recognition for Low-Resource Indian Languages using Multi-Task conformer
- Transformer-based end-to-end speech recognition with residual Gaussian-based self-attention
- Conv-Transformer Transducer: Low Latency, Low Frame Rate, Streamable End-to-End Speech Recognition
- When Can Self-Attention Be Replaced by Feed Forward Layers?
- End-to-End Dereverberation, Beamforming, and Speech Recognition with Improved Numerical Stability and Advanced Frontend
- Transferring Source Style in Non-Parallel Voice Conversion
- SkinAugment: Auto-Encoding Speaker Conversions for Automatic Speech Translation
- Tiny Transducer: A Highly-efficient Speech Recognition Model on Edge Devices
- Self-and-Mixed Attention Decoder with Deep Acoustic Structure for Transformer-based LVCSR
- Do as I mean, not as I say: Sequence Loss Training for Spoken Language Understanding
- End-to-End Automatic Speech Recognition Integrated With CTC-Based Voice Activity Detection
- Detecting Semantic Clones of Unseen Functionality
- Faster, Simpler and More Accurate Hybrid ASR Systems Using Wordpieces
- End-to-End Speaker-Attributed ASR with Transformer
- Domain Adaptation of NMT models for English-Hindi Machine Translation Task at AdapMT ICON 2020
- Attention-based ASR with Lightweight and Dynamic Convolutions
- Parallel Rescoring with Transformer for Streaming On-Device Speech Recognition
- Monolingual Data Selection Analysis for English-Mandarin Hybrid Code-switching Speech Recognition
- Streaming Transformer-based Acoustic Models Using Self-attention with Augmented Memory
- A Study of Non-autoregressive Model for Sequence Generation
- ADVISER: A Toolkit for Developing Multi-modal, Multi-domain and Socially-engaged Conversational Agents
- CAT: A CTC-CRF based ASR Toolkit Bridging the Hybrid and the End-to-end Approaches towards Data Efficiency and Low Latency
- Head-synchronous Decoding for Transformer-based Streaming ASR
- Label-Synchronous Speech-to-Text Alignment for ASR Using Forward and Backward Transformers
- Advanced Long-context End-to-end Speech Recognition Using Context-expanded Transformers
- Flexi-Transducer: Optimizing Latency, Accuracy and Compute forMulti-Domain On-Device Scenarios
- Internal Language Model Training for Domain-Adaptive End-to-End Speech Recognition
- Multi-QuartzNet: Multi-Resolution Convolution for Speech Recognition with Multi-Layer Feature Fusion
- Transformer in action: a comparative study of transformer-based acoustic models for large scale speech recognition applications
- Cross-attention conformer for context modeling in speech enhancement for ASR
- A Dual-Decoder Conformer for Multilingual Speech Recognition
- Streaming End-to-End ASR based on Blockwise Non-Autoregressive Models
- Recurrent Transformer-Based Near- and Far-Field THz Wideband Channel Estimation for UM-MIMO
- Bidirectional Representations for Low Resource Spoken Language Understanding
- Attention-based Multi-hypothesis Fusion for Speech Summarization
- Relaxing the Conditional Independence Assumption of CTC-based ASR by Conditioning on Intermediate Predictions
- Non-autoregressive Mandarin-English Code-switching Speech Recognition
- Effective Decoder Masking for Transformer Based End-to-End Speech Recognition
- Phoneme Based Neural Transducer for Large Vocabulary Speech Recognition
- Representation Learning for Sequence Data with Deep Autoencoding Predictive Components
- Streaming Transformer ASR with Blockwise Synchronous Beam Search
- Simplified Self-Attention for Transformer-based End-to-End Speech Recognition
- The RWTH ASR System for TED-LIUM Release 2: Improving Hybrid HMM with SpecAugment
- XPPG-PCA: Reference-free automatic speech severity evaluation with principal components
- Minimum Word Error Rate Training with Language Model Fusion for End-to-End Speech Recognition
- Semantic Data Augmentation for End-to-End Mandarin Speech Recognition
- Data Augmentation Methods for End-to-end Speech Recognition on Distant-Talk Scenarios
- Siamese Neural Networks for Class Activity Detection
- Multi-task Learning with Cross Attention for Keyword Spotting
- CASS-NAT: CTC Alignment-based Single Step Non-autoregressive Transformer for Speech Recognition
- An Improved Single Step Non-autoregressive Transformer for Automatic Speech Recognition
- Optimizing Latency for Online Video CaptioningUsing Audio-Visual Transformers
- Coarse-To-Fine And Cross-Lingual ASR Transfer
- Approaches to Improving Recognition of Underrepresented Named Entities in Hybrid ASR Systems
- Fast-MD: Fast Multi-Decoder End-to-End Speech Translation with Non-Autoregressive Hidden Intermediates
- Wav-BERT: Cooperative Acoustic and Linguistic Representation Learning for Low-Resource Speech Recognition
- An Investigation of Enhancing CTC Model for Triggered Attention-based Streaming ASR
- Cross-lingual Transfer for Speech Processing using Acoustic Language Similarity
- Far-Field Automatic Speech Recognition
- Improving RNN Transducer With Target Speaker Extraction and Neural Uncertainty Estimation
- Focus on the present: a regularization method for the ASR source-target attention layer
- Exploring Lexicon-Free Modeling Units for End-to-End Korean and Korean-English Code-Switching Speech Recognition
- Gender domain adaptation for automatic speech recognition task
- Transformer Based Deliberation for Two-Pass Speech Recognition
- Train your classifier first: Cascade Neural Networks Training from upper layers to lower layers
- Gaussian Kernelized Self-Attention for Long Sequence Data and Its Application to CTC-based Speech Recognition
- Integrating Knowledge into End-to-End Speech Recognition from External Text-Only Data
- Toward Streaming ASR with Non-Autoregressive Insertion-based Model
- Non-local convolutional neural networks (nlcnn) for speaker recognition
- Fast offline Transformer-based end-to-end automatic speech recognition for real-world applications
- Learning Speaker Embedding from Text-to-Speech
- Dual-Path Modeling for Long Recording Speech Separation in Meetings
- Cross-domain Speech Recognition with Unsupervised Character-level Distribution Matching