Achieving Human Parity in Conversational Speech Recognition
arXiv:1610.05256
Abstract
Conversational speech recognition has served as a flagship speech recognition task since the release of the Switchboard corpus in the 1990s. In this paper, we measure the human error rate on the widely used NIST 2000 test set, and find that our latest automated system has reached human parity. The error rate of professional transcribers is 5.9% for the Switchboard portion of the data, in which newly acquainted pairs of people discuss an assigned topic, and 11.3% for the CallHome portion where friends and family members have open-ended conversations. In both cases, our automated system establishes a new state of the art, and edges past the human benchmark, achieving error rates of 5.8% and 11.0%, respectively. The key to our system's performance is the use of various convolutional and LSTM acoustic model architectures, combined with a novel spatial smoothing method and lattice-free MMI acoustic training, multiple recurrent neural network language modeling approaches, and a systematic use of system combination.
Revised for publication, updated results
References in corpus (2)
Cited by in corpus (119)
- On Calibration of Modern Neural Networks
- DeepXplore: Automated Whitebox Testing of Deep Learning Systems
- Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning
- Cardiologist-Level Arrhythmia Detection with Convolutional Neural Networks
- Certified Defenses against Adversarial Examples
- The Microsoft 2017 Conversational Speech Recognition System
- SearchQA: A New Q&A Dataset Augmented with Context from a Search Engine
- The Microsoft 2016 Conversational Speech Recognition System
- A Simple Method for Commonsense Reasoning
- Snips Voice Platform: an embedded Spoken Language Understanding system for private-by-design voice interfaces
- Speaker-independent Speech Separation with Deep Attractor Network
- Cognitive Science in the era of Artificial Intelligence: A roadmap for reverse-engineering the infant language-learner
- Adversarial Sample Detection for Deep Neural Network through Model Mutation Testing
- How to Prove Your Model Belongs to You: A Blind-Watermark based Framework to Protect Intellectual Property of DNN
- Review: Deep Learning in Electron Microscopy
- Trustless Machine Learning Contracts; Evaluating and Exchanging Machine Learning Models on the Ethereum Blockchain
- Recognizing Overlapped Speech in Meetings: A Multichannel Separation Approach Using Neural Networks
- Factorization tricks for LSTM networks
- Big-Little Net: An Efficient Multi-Scale Feature Representation for Visual and Speech Recognition
- Adversarial Example Defenses: Ensembles of Weak Defenses are not Strong
- The CAPIO 2017 Conversational Speech Recognition System
- Exploring Neural Transducers for End-to-End Speech Recognition
- Comparing Human and Machine Errors in Conversational Speech Transcription
- Transformer-Transducer: End-to-End Speech Recognition with Self-Attention
- Advancing the State of the Art in Open Domain Dialog Systems through the Alexa Prize
- Learning Combinations of Activation Functions
- A Survey of Adversarial Learning on Graphs
- Evaluating the Usability of Automatically Generated Captions for People who are Deaf or Hard of Hearing
- Gram-CTC: Automatic Unit Selection and Target Decomposition for Sequence Labelling
- Deep learning for clustering of continuous gravitational wave candidates
- Deep Learning for Computational Chemistry
- Deep multi-survey classification of variable stars
- Dense Prediction on Sequences with Time-Dilated Convolutions for Speech Recognition
- The Airbus Air Traffic Control speech recognition 2018 challenge: towards ATC automatic transcription and call sign detection
- Adversarial Examples - A Complete Characterisation of the Phenomenon
- Improving Electron Micrograph Signal-to-Noise with an Atrous Convolutional Encoder-Decoder
- UniSpeech: Unified Speech Representation Learning with Labeled and Unlabeled Data
- A Method for Analysis of Patient Speech in Dialogue for Dementia Detection
- TinySpeech: Attention Condensers for Deep Speech Recognition Neural Networks on Edge Devices
- Leveraging native language information for improved accented speech recognition
- Confusion2Vec: Towards Enriching Vector Space Word Representations with Representational Ambiguities
- VoiceMask: Anonymize and Sanitize Voice Input on Mobile Devices
- Concept Learning through Deep Reinforcement Learning with Memory-Augmented Neural Networks
- Lattice Long Short-Term Memory for Human Action Recognition
- Deep-Dup: An Adversarial Weight Duplication Attack Framework to Crush Deep Neural Network in Multi-Tenant FPGA
- DeepHammer: Depleting the Intelligence of Deep Neural Networks through Targeted Chain of Bit Flips
- Estimate and Replace: A Novel Approach to Integrating Deep Neural Networks with Existing Applications
- Advances in Online Audio-Visual Meeting Transcription
- A Comparison of Online Automatic Speech Recognition Systems and the Nonverbal Responses to Unintelligible Speech
- The History of Speech Recognition to the Year 2030
- Unsupervised Pretraining for Sequence to Sequence Learning
- How Much Can We Really Trust You? Towards Simple, Interpretable Trust Quantification Metrics for Deep Neural Networks
- Meeting Transcription Using Virtual Microphone Arrays
- Accented Speech Recognition: A Survey
- Feature Extraction for Temporal Signal Recognition: An Overview
- Gradient Band-based Adversarial Training for Generalized Attack Immunity of A3C Path Finding
- Guided Source Separation Meets a Strong ASR Backend: Hitachi/Paderborn University Joint Investigation for Dinner Party ASR
- Recognizing Multi-talker Speech with Permutation Invariant Training
- Neural network gradient-based learning of black-box function interfaces
- Multitask Learning with CTC and Segmental CRF for Speech Recognition
- DeepGini: Prioritizing Massive Tests to Enhance the Robustness of Deep Neural Networks
- Unsupervised Domain Adaptation by Adversarial Learning for Robust Speech Recognition
- ATHENA: A Framework based on Diverse Weak Defenses for Building Adversarial Defense
- Reducing Bias in Production Speech Models
- Automatic speech recognition for the Nepali language using CNN, bidirectional LSTM and ResNet
- Blind Pre-Processing: A Robust Defense Method Against Adversarial Examples
- Denoised Internal Models: a Brain-Inspired Autoencoder against Adversarial Attacks
- Enhancing Supermarket Robot Interaction: A Multi-Level LLM Conversational Interface for Handling Diverse Customer Intents
- Multi-Class Gaussian Process Classification Made Conjugate: Efficient Inference via Data Augmentation
- Detecting Adversarial Examples via Neural Fingerprinting
- Comparison-Based Convolutional Neural Networks for Cervical Cell/Clumps Detection in the Limited Data Scenario
- Small-Footprint Open-Vocabulary Keyword Spotting with Quantized LSTM Networks
- Nonsense Attacks on Google Assistant
- Simplified End-to-End MMI Training and Voting for ASR
- Joint Separation and Denoising of Noisy Multi-talker Speech using Recurrent Neural Networks and Permutation Invariant Training
- Joint Modeling of Accents and Acoustics for Multi-Accent Speech Recognition
- Dialectal Speech Recognition and Translation of Swiss German Speech to Standard German Text: Microsoft's Submission to SwissText 2021
- A comprehensive study of batch construction strategies for recurrent neural networks in MXNet
- Cross-utterance Reranking Models with BERT and Graph Convolutional Networks for Conversational Speech Recognition
- Cross-Utterance Language Models with Acoustic Error Sampling
- Capacity Control of ReLU Neural Networks by Basis-path Norm
- On Feature Decorrelation in Self-Supervised Learning
- Phonemic and Graphemic Multilingual CTC Based Speech Recognition
- Transformer Language Models with LSTM-based Cross-utterance Information Representation
- Compressing LSTM Networks by Matrix Product Operators
- ASAPP-ASR: Multistream CNN and Self-Attentive SRU for SOTA Speech Recognition
- Neural Task Representations as Weak Supervision for Model Agnostic Cross-Lingual Transfer
- Recent Progresses in Deep Learning based Acoustic Models (Updated)
- VoiceMoji: A Novel On-Device Pipeline for Seamless Emoji Insertion in Dictation
- Auxiliary Interference Speaker Loss for Target-Speaker Speech Recognition
- AIx Speed: Playback Speed Optimization Using Listening Comprehension of Speech Recognition Models
- Analyzing deep CNN-based utterance embeddings for acoustic model adaptation
- Dompteur: Taming Audio Adversarial Examples
- Combining Spatial Clustering with LSTM Speech Models for Multichannel Speech Enhancement
- Towards Deep Learning Models Resistant to Large Perturbations
- Unsupervised Speech Recognition via Segmental Empirical Output Distribution Matching
- When CTC Training Meets Acoustic Landmarks
- Tag and correct: high precision post-editing approach to correction of speech recognition errors
- Language Modeling with Highway LSTM
- QASR: QCRI Aljazeera Speech Resource -- A Large Scale Annotated Arabic Speech Corpus
- Improved MVDR Beamforming Using LSTM Speech Models to Clean Spatial Clustering Masks
- Cascaded CNN-resBiLSTM-CTC: An End-to-End Acoustic Model For Speech Recognition
- Defending Against Adversarial Denial-of-Service Data Poisoning Attacks
- Text-based classification of interviews for mental health -- juxtaposing the state of the art
- Multi-Talker MVDR Beamforming Based on Extended Complex Gaussian Mixture Model
- Classification of sparsely labeled spatio-temporal data through semi-supervised adversarial learning
- Utterance-level Permutation Invariant Training with Latency-controlled BLSTM for Single-channel Multi-talker Speech Separation
- A Comparison of Lattice-free Discriminative Training Criteria for Purely Sequence-Trained Neural Network Acoustic Models
- (Pen-) Ultimate DNN Pruning
- Look-up and Adapt: A One-shot Semantic Parser
- Transforming unstructured voice and text data into insight for paramedic emergency service using recurrent and convolutional neural networks
- Automatic, Dynamic, and Nearly Optimal Learning Rate Specification by Local Quadratic Approximation
- Enhancement of Spatial Clustering-Based Time-Frequency Masks using LSTM Neural Networks
- Exploring End-to-End Multi-channel ASR with Bias Information for Meeting Transcription
- Graph-Based Fuzz Testing for Deep Learning Inference Engine
- SANTLR: Speech Annotation Toolkit for Low Resource Languages
- Multi-Frame Cross-Entropy Training for Convolutional Neural Networks in Speech Recognition
- Akid: A Library for Neural Network Research and Production from a Dataism Approach
- Multi-talker Speech Separation with Utterance-level Permutation Invariant Training of Deep Recurrent Neural Networks