Deep Speech 2: End-to-End Speech Recognition in English and Mandarin
arXiv:1512.02595
Abstract
We show that an end-to-end deep learning approach can be used to recognize either English or Mandarin Chinese speech--two vastly different languages. Because it replaces entire pipelines of hand-engineered components with neural networks, end-to-end learning allows us to handle a diverse variety of speech including noisy environments, accents and different languages. Key to our approach is our application of HPC techniques, resulting in a 7x speedup over our previous system. Because of this efficiency, experiments that previously took weeks now run in days. This enables us to iterate more quickly to identify superior architectures and algorithms. As a result, in several cases, our system is competitive with the transcription of human workers when benchmarked on standard datasets. Finally, using a technique called Batch Dispatch with GPUs in the data center, we show that our system can be inexpensively deployed in an online setting, delivering low latency when serving users at scale.
References in corpus (6)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- On the difficulty of training Recurrent Neural Networks
- Going Deeper with Convolutions
- cuDNN: Efficient Primitives for Deep Learning
- End-to-end Continuous Speech Recognition using Attention-based Recurrent NN: First Results
- First-Pass Large Vocabulary Continuous Speech Recognition using Bi-Directional Recurrent DNNs
Cited by in corpus (511)
- mixup: Beyond Empirical Risk Minimization
- xDeepFM: Combining Explicit and Implicit Feature Interactions for Recommender Systems
- Mixed Precision Training
- A Survey on Distributed Machine Learning
- A Convergence Theory for Deep Learning via Over-Parameterization
- Convolutional Recurrent Neural Networks for Polyphonic Sound Event Detection
- Cardiologist-Level Arrhythmia Detection with Convolutional Neural Networks
- Achieving Human Parity in Conversational Speech Recognition
- Self-Supervised Representation Learning: Introduction, Advances and Challenges
- Deep Autoencoder based Energy Method for the Bending, Vibration, and Buckling Analysis of Kirchhoff Plates
- Deep Speaker: an End-to-End Neural Speaker Embedding System
- Light Gated Recurrent Units for Speech Recognition
- Train longer, generalize better: closing the generalization gap in large batch training of neural networks
- Deep Voice: Real-time Neural Text-to-Speech
- Neural Network Detection of Data Sequences in Communication Systems
- Recent Advances in Convolutional Neural Networks
- Fast Parallel Hypertree Decompositions in Logarithmic Recursion Depth
- Compute Trends Across Three Eras of Machine Learning
- Improved training of end-to-end attention models for speech recognition
- XSleepNet: Multi-View Sequential Model for Automatic Sleep Staging
- TED-LIUM 3: twice as much data and corpus repartition for experiments on speaker adaptation
- Wav2Letter: an End-to-End ConvNet-based Speech Recognition System
- Monitoring COVID-19 social distancing with person detection and tracking via fine-tuned YOLO v3 and Deepsort techniques
- Benchmarking TPU, GPU, and CPU Platforms for Deep Learning
- A Microprocessor implemented in 65nm CMOS with Configurable and Bit-scalable Accelerator for Programmable In-memory Computing
- Realtime Robust Malicious Traffic Detection via Frequency Domain Analysis
- Learning to Predict the Cosmological Structure Formation
- Cognitive Science in the era of Artificial Intelligence: A roadmap for reverse-engineering the infant language-learner
- Quantum Entanglement in Deep Learning Architectures
- Demystifying Parallel and Distributed Deep Learning: An In-Depth Concurrency Analysis
- Pushing the Limits of Semi-Supervised Learning for Automatic Speech Recognition
- A Survey on Methods and Theories of Quantized Neural Networks
- Did you hear that? Adversarial Examples Against Automatic Speech Recognition
- Natural Language Processing Advancements By Deep Learning: A Survey
- SCALE-Sim: Systolic CNN Accelerator Simulator
- LipNet: End-to-End Sentence-level Lipreading
- Comparing Speech and Keyboard Text Entry for Short Messages in Two Languages on Touchscreen Phones
- Fathom: Reference Workloads for Modern Deep Learning Methods
- Neural Voice Cloning with a Few Samples
- L1-Norm Batch Normalization for Efficient Training of Deep Neural Networks
- Robust Audio Adversarial Example for a Physical Attack
- Threat of Adversarial Attacks on Deep Learning in Computer Vision: A Survey
- Houdini: Fooling Deep Structured Prediction Models
- End-to-end ASR: from Supervised to Semi-Supervised Learning with Modern Architectures
- Deep learning incorporating biologically-inspired neural dynamics
- Misalignment Resilient Diffractive Optical Networks
- CNN+LSTM Architecture for Speech Emotion Recognition with Data Augmentation
- Curriculum Learning by Transfer Learning: Theory and Experiments with Deep Networks
- Evaluating Bayesian Deep Learning Methods for Semantic Segmentation
- Beyond Data and Model Parallelism for Deep Neural Networks
- DSD: Dense-Sparse-Dense Training for Deep Neural Networks
- Bayesian Recurrent Neural Networks
- A Survey of FPGA-Based Neural Network Accelerator
- Spinal cord gray matter segmentation using deep dilated convolutions
- Exploring Sparsity in Recurrent Neural Networks
- Self-Attention Networks for Connectionist Temporal Classification in Speech Recognition
- Towards Automatic Face-to-Face Translation
- Emotion Recognition From Speech With Recurrent Neural Networks
- Deep Learning: A Bayesian Perspective
- A neural attention model for speech command recognition
- A Survey on Non-Intrusive Load Monitoring Methodies and Techniques for Energy Disaggregation Problem
- Recent Advances in Deep Learning: An Overview
- Compressing Recurrent Neural Network with Tensor Train
- On the Convergence Rate of Training Recurrent Neural Networks
- Hardware Acceleration of Sparse and Irregular Tensor Computations of ML Models: A Survey and Insights
- Dataset Condensation with Gradient Matching
- High Fidelity Speech Synthesis with Adversarial Networks
- Deceiving End-to-End Deep Learning Malware Detectors using Adversarial Examples
- Deep Learning for Android Malware Defenses: a Systematic Literature Review
- Recognizing Overlapped Speech in Meetings: A Multichannel Separation Approach Using Neural Networks
- Bottom-up and top-down approaches for the design of neuromorphic processing systems: Tradeoffs and synergies between natural and artificial intelligence
- Block-Sparse Recurrent Neural Networks
- Communication-Efficient Distributed Deep Learning: A Comprehensive Survey
- Distribution Aligning Refinery of Pseudo-label for Imbalanced Semi-supervised Learning
- STRIP: A Defence Against Trojan Attacks on Deep Neural Networks
- Character-Level Question Answering with Attention
- Self-Attention Transducers for End-to-End Speech Recognition
- FEA-Net: A Physics-guided Data-driven Model for Efficient Mechanical Response Prediction
- Fully Convolutional Speech Recognition
- Optimizing Performance of Recurrent Neural Networks on GPUs
- Learning without feedback: Fixed random learning signals allow for feedforward training of deep neural networks
- CommanderSong: A Systematic Approach for Practical Adversarial Voice Recognition
- Unsupervised Automatic Speech Recognition: A Review
- ESPnet: End-to-End Speech Processing Toolkit
- FaceFilter: Audio-visual speech separation using still images
- Exploring Neural Transducers for End-to-End Speech Recognition
- Scaling Deep Learning on GPU and Knights Landing clusters
- NSML: A Machine Learning Platform That Enables You to Focus on Your Models
- Fully Distributed Multi-Robot Collision Avoidance via Deep Reinforcement Learning for Safe and Efficient Navigation in Complex Scenarios
- Task Agnostic Continual Learning Using Online Variational Bayes
- Multi-Modal Emotion recognition on IEMOCAP Dataset using Deep Learning
- Do Explanations Reflect Decisions? A Machine-centric Strategy to Quantify the Performance of Explainability Algorithms
- Residual Convolutional CTC Networks for Automatic Speech Recognition
- Deep learning as a tool for neural data analysis: speech classification and cross-frequency coupling in human sensorimotor cortex
- A Study of BFLOAT16 for Deep Learning Training
- Online Hybrid CTC/Attention End-to-End Automatic Speech Recognition Architecture
- Improved Recurrent Neural Networks for Session-based Recommendations
- Towards Relatable Explainable AI with the Perceptual Process
- Deep Learning for Environmentally Robust Speech Recognition: An Overview of Recent Developments
- ESE: Efficient Speech Recognition Engine with Sparse LSTM on FPGA
- Machine learning for music genre: multifaceted review and experimentation with audioset
- Towards better decoding and language model integration in sequence to sequence models
- Defining the scope of AI regulations
- Neural Speech Recognizer: Acoustic-to-Word LSTM Model for Large Vocabulary Speech Recognition
- Gram-CTC: Automatic Unit Selection and Target Decomposition for Sequence Labelling
- TBD: Benchmarking and Analyzing Deep Neural Network Training
- Real-time Neural Network Inference on Extremely Weak Devices: Agile Offloading with Explainable AI
- Recurrent Batch Normalization
- i-Mix: A Domain-Agnostic Strategy for Contrastive Representation Learning
- Good Subnetworks Provably Exist: Pruning via Greedy Forward Selection
- SF-Net: Structured Feature Network for Continuous Sign Language Recognition
- Spartus: A 9.4 TOp/s FPGA-based LSTM Accelerator Exploiting Spatio-Temporal Sparsity
- Recurrent Neural Networks With Limited Numerical Precision
- Quadratic Suffices for Over-parametrization via Matrix Chernoff Bound
- HRel: Filter Pruning based on High Relevance between Activation Maps and Class Labels
- Active Learning for Speech Recognition: the Power of Gradients
- Towards Training Reproducible Deep Learning Models
- Unified Streaming and Non-streaming Two-pass End-to-end Model for Speech Recognition
- ImageNet Training in Minutes
- Scaling data-driven robotics with reward sketching and batch reinforcement learning
- A 23 W Keyword Spotting IC with Ring-Oscillator-Based Time-Domain Feature Extraction
- Jasper: An End-to-End Convolutional Neural Acoustic Model
- Generative Language Modeling for Automated Theorem Proving
- Towards End-to-End Code-Switching Speech Recognition
- Citrinet: Closing the Gap between Non-Autoregressive and Autoregressive End-to-End Models for Automatic Speech Recognition
- Delta Networks for Optimized Recurrent Network Computation
- Direct Speech-to-image Translation
- Mixed-Precision Training for NLP and Speech Recognition with OpenSeq2Seq
- Robust Training of Vector Quantized Bottleneck Models
- Training Neural Speech Recognition Systems with Synthetic Speech Augmentation
- Transformer-based End-to-End Speech Recognition with Local Dense Synthesizer Attention
- A Highly Adaptive Acoustic Model for Accurate Multi-Dialect Speech Recognition
- ARM-Net: Adaptive Relation Modeling Network for Structured Data
- Successes and critical failures of neural networks in capturing human-like speech recognition
- Sequence Modeling via Segmentations
- Hierarchical Multitask Learning for CTC-based Speech Recognition
- Fooling OCR Systems with Adversarial Text Images
- End-to-end Audiovisual Speech Activity Detection with Bimodal Recurrent Neural Models
- A Noise-Robust Self-supervised Pre-training Model Based Speech Representation Learning for Automatic Speech Recognition
- Online Normalization for Training Neural Networks
- Energy-based error bound of physics-informed neural network solutions in elasticity
- Letter-Based Speech Recognition with Gated ConvNets
- Online Keyword Spotting with a Character-Level Recurrent Neural Network
- Whispering LLaMA: A Cross-Modal Generative Error Correction Framework for Speech Recognition
- ATCSpeech: a multilingual pilot-controller speech corpus from real Air Traffic Control environment
- An End-to-End Architecture for Keyword Spotting and Voice Activity Detection
- AIBench: An Industry Standard Internet Service AI Benchmark Suite
- Towards Online End-to-end Transformer Automatic Speech Recognition
- Learning Multiscale Features Directly From Waveforms
- ElasticTrainer: Speeding Up On-Device Training with Runtime Elastic Tensor Selection
- Stochastic Mirror Descent on Overparameterized Nonlinear Models: Convergence, Implicit Regularization, and Generalization
- Reusing Neural Speech Representations for Auditory Emotion Recognition
- Opening the black box of neural nets: case studies in stop/top discrimination
- Discrete and continuous representations and processing in deep learning: Looking forward
- A Tutorial on Ultra-Reliable and Low-Latency Communications in 6G: Integrating Domain Knowledge into Deep Learning
- Who Needs Words? Lexicon-Free Speech Recognition
- Leveraging native language information for improved accented speech recognition
- Toward a Better Monitoring Statistic for Profile Monitoring via Variational Autoencoders
- Opportunities and Challenges in Deep Learning Adversarial Robustness: A Survey
- Semi-Supervised Speech Recognition via Local Prior Matching
- Throughput Optimizations for FPGA-based Deep Neural Network Inference
- A Full Text-Dependent End to End Mispronunciation Detection and Diagnosis with Easy Data Augmentation Techniques
- One Size Does Not Fit All: Quantifying and Exposing the Accuracy-Latency Trade-off in Machine Learning Cloud Service APIs via Tolerance Tiers
- Unsupervised State Representation Learning in Atari
- Heterogeneity-Aware Cluster Scheduling Policies for Deep Learning Workloads
- Unpaired Speech Enhancement by Acoustic and Adversarial Supervision for Speech Recognition
- Improving the Neural GPU Architecture for Algorithm Learning
- Pollux: Co-adaptive Cluster Scheduling for Goodput-Optimized Deep Learning
- Parallelizing Linear Recurrent Neural Nets Over Sequence Length
- Named Entity Recognition with stack residual LSTM and trainable bias decoding
- Predicting Auction Price of Vehicle License Plate with Deep Recurrent Neural Network
- Universal adversarial examples in speech command classification
- Streaming Normalization: Towards Simpler and More Biologically-plausible Normalizations for Online and Recurrent Learning
- CrossASR++: A Modular Differential Testing Framework for Automatic Speech Recognition
- Improving Generalization of Transformer for Speech Recognition with Parallel Schedule Sampling and Relative Positional Embedding
- Low Latency End-to-End Streaming Speech Recognition with a Scout Network
- End-to-End Deep Fault Tolerant Control
- U2++: Unified Two-pass Bidirectional End-to-end Model for Speech Recognition
- Automated Question Answer medical model based on Deep Learning Technology
- Unsupervised Speech Recognition
- Sequence Discriminative Training for Deep Learning based Acoustic Keyword Spotting
- Advances in Joint CTC-Attention based End-to-End Speech Recognition with a Deep CNN Encoder and RNN-LM
- No Padding Please: Efficient Neural Handwriting Recognition
- BioNetExplorer: Architecture-Space Exploration of Bio-Signal Processing Deep Neural Networks for Wearables
- Handcrafted Backdoors in Deep Neural Networks
- A Spectral Energy Distance for Parallel Speech Synthesis
- Highrisk Prediction from Electronic Medical Records via Deep Attention Networks
- Escoin: Efficient Sparse Convolutional Neural Network Inference on GPUs
- Convo: What does conversational programming need? An exploration of machine learning interface design
- Multistage linguistic conditioning of convolutional layers for speech emotion recognition
- End-to-end named entity extraction from speech
- Deep Learning at 15PF: Supervised and Semi-Supervised Classification for Scientific Data
- Preech: A System for Privacy-Preserving Speech Transcription
- A Simplified Fully Quantized Transformer for End-to-end Speech Recognition
- Low-frequency Compensated Synthetic Impulse Responses for Improved Far-field Speech Recognition
- Convolutional RNN: an Enhanced Model for Extracting Features from Sequential Data
- Deep Recurrent Convolutional Neural Network: Improving Performance For Speech Recognition
- Multilingual Graphemic Hybrid ASR with Massive Data Augmentation
- Pushing the Limits of Non-Autoregressive Speech Recognition
- An Investigation of End-to-End Multichannel Speech Recognition for Reverberant and Mismatch Conditions
- Removing Backdoor-Based Watermarks in Neural Networks with Limited Data
- QuClassi: A Hybrid Deep Neural Network Architecture based on Quantum State Fidelity
- Imperceptible Adversarial Examples by Spatial Chroma-Shift
- The History of Speech Recognition to the Year 2030
- WeNet: Production oriented Streaming and Non-streaming End-to-End Speech Recognition Toolkit
- 4K-Memristor Analog-Grade Passive Crossbar Circuit
- A holistic approach to polyphonic music transcription with neural networks
- Multi-Dialect Arabic Speech Recognition
- DeepIntent: ImplicitIntent based Android IDS with E2E Deep Learning architecture
- Understanding Reuse, Performance, and Hardware Cost of DNN Dataflows: A Data-Centric Approach Using MAESTRO
- LSTMVis: A Tool for Visual Analysis of Hidden State Dynamics in Recurrent Neural Networks
- A Fully Differentiable Beam Search Decoder
- Robust Watermarking of Neural Network with Exponential Weighting
- Rethinking Full Connectivity in Recurrent Neural Networks
- Image denoising and restoration with CNN-LSTM Encoder Decoder with Direct Attention
- Exploring the Design Space of Deep Convolutional Neural Networks at Large Scale
- Espresso: A Fast End-to-end Neural Speech Recognition Toolkit
- Lip-to-Speech Synthesis for Arbitrary Speakers in the Wild
- Neuropathic Pain Diagnosis Simulator for Causal Discovery Algorithm Evaluation
- Speech-Based Visual Question Answering
- The Architectural Implications of Facebook's DNN-based Personalized Recommendation
- SparseTrain:Leveraging Dynamic Sparsity in Training DNNs on General-Purpose SIMD Processors
- Functions that Emerge through End-to-End Reinforcement Learning - The Direction for Artificial General Intelligence -
- Luck Matters: Understanding Training Dynamics of Deep ReLU Networks
- Focal Loss based Residual Convolutional Neural Network for Speech Emotion Recognition
- Real-Time Machine Learning: The Missing Pieces
- Training Robust Deep Neural Networks via Adversarial Noise Propagation
- Orchestrating the Development Lifecycle of Machine Learning-Based IoT Applications: A Taxonomy and Survey
- Reinforcement Learning and Adaptive Sampling for Optimized DNN Compilation
- Recent Advances in End-to-End Spoken Language Understanding
- A Study of Deep Learning Robustness Against Computation Failures
- Bayesian Sparsification of Recurrent Neural Networks
- Learning Robust and Multilingual Speech Representations
- Semi-Supervised Model Training for Unbounded Conversational Speech Recognition
- AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head Synthesis
- Deep-Energy: Unsupervised Training of Deep Neural Networks
- End-to-End Spoken Language Understanding: Performance analyses of a voice command task in a low resource setting
- Emphasizing Unseen Words: New Vocabulary Acquisition for End-to-End Speech Recognition
- Dialogue history integration into end-to-end signal-to-concept spoken language understanding systems
- Towards End-to-end Automatic Code-Switching Speech Recognition
- Guided Source Separation Meets a Strong ASR Backend: Hitachi/Paderborn University Joint Investigation for Dinner Party ASR
- Learning Fast and Slow: PROPEDEUTICA for Real-time Malware Detection
- A novel pyramidal-FSMN architecture with lattice-free MMI for speech recognition
- Demystifying Differentiable Programming: Shift/Reset the Penultimate Backpropagator
- Improved Regularization Techniques for End-to-End Speech Recognition
- RNN-T Models Fail to Generalize to Out-of-Domain Audio: Causes and Solutions
- Recognizing Multi-talker Speech with Permutation Invariant Training
- Deep-Learned Collision Avoidance Policy for Distributed Multi-Agent Navigation
- FALCON: A Fourier Transform Based Approach for Fast and Secure Convolutional Neural Network Predictions
- An improved hybrid CTC-Attention model for speech recognition
- Learning spectro-temporal representations of complex sounds with parameterized neural networks
- Efficient conformer-based speech recognition with linear attention
- An investigation of phone-based subword units for end-to-end speech recognition
- Unbounded cache model for online language modeling with open vocabulary
- An Empirical Study of Large-Batch Stochastic Gradient Descent with Structured Covariance Noise
- Neuroevolution of Neural Network Architectures Using CoDeepNEAT and Keras
- On Arrhythmia Detection by Deep Learning and Multidimensional Representation
- CAT: CRF-based ASR Toolkit
- Adversarial Attacks and Defenses for Speech Recognition Systems
- Speaker-Conditioned Hierarchical Modeling for Automated Speech Scoring
- Straggler-Resilient Distributed Machine Learning with Dynamic Backup Workers
- Training speaker recognition systems with limited data
- WaveGuard: Understanding and Mitigating Audio Adversarial Examples
- WNARS: WFST based Non-autoregressive Streaming End-to-End Speech Recognition
- FrugalML: How to Use ML Prediction APIs More Accurately and Cheaply
- The Case for Strong Scaling in Deep Learning: Training Large 3D CNNs with Hybrid Parallelism
- From Senones to Chenones: Tied Context-Dependent Graphemes for Hybrid Speech Recognition
- A SOT-MRAM-based Processing-In-Memory Engine for Highly Compressed DNN Implementation
- Improved training for online end-to-end speech recognition systems
- Improving LSTM-CTC based ASR performance in domains with limited training data
- A Hardware-Oriented and Memory-Efficient Method for CTC Decoding
- Do End-to-End Speech Recognition Models Care About Context?
- Factorized Fourier Neural Operators
- A Multi-Task Learning Framework for Overcoming the Catastrophic Forgetting in Automatic Speech Recognition
- Attention-Augmented End-to-End Multi-Task Learning for Emotion Prediction from Speech
- Denoised Internal Models: a Brain-Inspired Autoencoder against Adversarial Attacks
- Incremental Learning Using a Grow-and-Prune Paradigm with Efficient Neural Networks
- Leveraging End-to-End Speech Recognition with Neural Architecture Search
- Revisiting BFloat16 Training
- Detection based Defense against Adversarial Examples from the Steganalysis Point of View
- Escaping Saddle Points Faster with Stochastic Momentum
- Exploiting Spectral Augmentation for Code-Switched Spoken Language Identification
- A Deep-learning-based Method for PIR-based Multi-person Localization
- Dance Dance Convolution
- SGAD: Soft-Guided Adaptively-Dropped Neural Network
- To Compress, or Not to Compress: Characterizing Deep Learning Model Compression for Embedded Inference
- Human-Machine Interaction Speech Corpus from the ROBIN project
- Machine Learning for Robust Identification of Complex Nonlinear Dynamical Systems: Applications to Earth Systems Modeling
- Improving End-to-End Speech Recognition with Policy Learning
- Beyond Human-Level Accuracy: Computational Challenges in Deep Learning
- Graphcore C2 Card performance for image-based deep learning application: A Report
- A Novel Fusion of Attention and Sequence to Sequence Autoencoders to Predict Sleepiness From Speech
- ClovaCall: Korean Goal-Oriented Dialog Speech Corpus for Automatic Speech Recognition of Contact Centers
- Beyond clipping: Equalization-based Psychoacoustic Attacks against ASRs
- Multi-modal Automated Speech Scoring using Attention Fusion
- Masked Pre-trained Encoder base on Joint CTC-Transformer
- Echo: Compiler-based GPU Memory Footprint Reduction for LSTM RNN Training
- Training Deep Neural Networks Using Posit Number System
- Relative stability toward diffeomorphisms indicates performance in deep nets
- Benchmarking the Performance and Energy Efficiency of AI Accelerators for AI Training
- Measuring the Effectiveness of Voice Conversion on Speaker Identification and Automatic Speech Recognition Systems
- Hard Sample Mining for the Improved Retraining of Automatic Speech Recognition
- EfficientWord-Net: An Open Source Hotword Detection Engine based on One-shot Learning
- Improving Noise Robustness of an End-to-End Neural Model for Automatic Speech Recognition
- Adaptive Dense-to-Sparse Paradigm for Pruning Online Recommendation System with Non-Stationary Data
- Exploring spectro-temporal features in end-to-end convolutional neural networks
- Ed-Fed: A generic federated learning framework with resource-aware client selection for edge devices
- Assessing SATNet's Ability to Solve the Symbol Grounding Problem
- Edge AIBench: Towards Comprehensive End-to-end Edge Computing Benchmarking
- Independent language modeling architecture for end-to-end ASR
- Insights on Neural Representations for End-to-End Speech Recognition
- A Multiversion Programming Inspired Approach to Detecting Audio Adversarial Examples
- SEC4SR: A Security Analysis Platform for Speaker Recognition
- An Online Attention-based Model for Speech Recognition
- DARTS: Dialectal Arabic Transcription System
- Auxiliary Multimodal LSTM for Audio-visual Speech Recognition and Lipreading
- FPGA-Based Low-Power Speech Recognition with Recurrent Neural Networks
- Audio-Linguistic Embeddings for Spoken Sentences
- Pansori: ASR Corpus Generation from Open Online Video Contents
- Phoenix: A Low-Precision Floating-Point Quantization Oriented Architecture for Convolutional Neural Networks
- XCloud: Design and Implementation of AI Cloud Platform with RESTful API Service
- Improving RNN Transducer Modeling for End-to-End Speech Recognition
- Estimating the Brittleness of AI: Safety Integrity Levels and the Need for Testing Out-Of-Distribution Performance
- Effects of Number of Filters of Convolutional Layers on Speech Recognition Model Accuracy
- End-to-end Jordanian dialect speech-to-text self-supervised learning framework
- Romanian Speech Recognition Experiments from the ROBIN Project
- A systematic comparison of grapheme-based vs. phoneme-based label units for encoder-decoder-attention models
- QuTiBench: Benchmarking Neural Networks on Heterogeneous Hardware
- Democratizing Production-Scale Distributed Deep Learning
- Gated Recurrent Unit Based Acoustic Modeling with Future Context
- A comparable study of modeling units for end-to-end Mandarin speech recognition
- Adaptive Selection of Deep Learning Models on Embedded Systems
- Joint Modeling of Accents and Acoustics for Multi-Accent Speech Recognition
- Research on Modeling Units of Transformer Transducer for Mandarin Speech Recognition
- Efficient Synthesis of Compact Deep Neural Networks
- A Generalizable Approach to Learning Optimizers
- Boosting Active Learning for Speech Recognition with Noisy Pseudo-labeled Samples
- Elastic Gossip: Distributing Neural Network Training Using Gossip-like Protocols
- Learning from Learning Machines: Optimisation, Rules, and Social Norms
- Privacy Inference Attacks and Defenses in Cloud-based Deep Neural Network: A Survey
- Evaluating Large Language Models in Code Generation: INFINITE Methodology for Defining the Inference Index
- On the Inductive Bias of Word-Character-Level Multi-Task Learning for Speech Recognition
- Audio Enhancement for Computer Audition -- An Iterative Training Paradigm Using Sample Importance
- Advancing CTC-CRF Based End-to-End Speech Recognition with Wordpieces and Conformers
- Perspective: A Phase Diagram for Deep Learning unifying Jamming, Feature Learning and Lazy Training
- A Comparison of Hybrid and End-to-End Models for Syllable Recognition
- WaveTTS: Tacotron-based TTS with Joint Time-Frequency Domain Loss
- Snore-GANs: Improving Automatic Snore Sound Classification with Synthesized Data
- A.I. based Embedded Speech to Text Using Deepspeech
- Phonemic and Graphemic Multilingual CTC Based Speech Recognition
- Transformer with Bidirectional Decoder for Speech Recognition
- Hybrid CTC-Attention based End-to-End Speech Recognition using Subword Units
- Learning Noise-Invariant Representations for Robust Speech Recognition
- A Survey of Techniques All Classifiers Can Learn from Deep Networks: Models, Optimizations, and Regularization
- A Simple yet Effective Baseline for Robust Deep Learning with Noisy Labels
- Exponential Moving Average Model in Parallel Speech Recognition Training
- Insertion-Based Modeling for End-to-End Automatic Speech Recognition
- Modelling of daily reference evapotranspiration using deep neural network in different climates
- Exploiting Contextual Information with Deep Neural Networks
- To Pretrain or Not to Pretrain: Examining the Benefits of Pretraining on Resource Rich Tasks
- Resource aware design of a deep convolutional-recurrent neural network for speech recognition through audio-visual sensor fusion
- deepSELF: An Open Source Deep Self End-to-End Learning Framework
- Complex Transformer: A Framework for Modeling Complex-Valued Sequence
- Recent Progresses in Deep Learning based Acoustic Models (Updated)
- AIBench Scenario: Scenario-distilling AI Benchmarking
- Modeling the Second Player in Distributionally Robust Optimization
- Multi-task Recurrent Model for True Multilingual Speech Recognition
- Memory Visualization for Gated Recurrent Neural Networks in Speech Recognition
- M3D-GAN: Multi-Modal Multi-Domain Translation with Universal Attention
- ReconVAT: A Semi-Supervised Automatic Music Transcription Framework for Low-Resource Real-World Data
- SparseNN: An Energy-Efficient Neural Network Accelerator Exploiting Input and Output Sparsity
- AIBench Training: Balanced Industry-Standard AI Training Benchmarking
- Auxiliary Interference Speaker Loss for Target-Speaker Speech Recognition
- A Sustainable Multi-modal Multi-layer Emotion-aware Service at the Edge
- Nonlinear Regression with a Convolutional Encoder-Decoder for Remote Monitoring of Surface Electrocardiograms
- End to End ASR System with Automatic Punctuation Insertion
- PDAugment: Data Augmentation by Pitch and Duration Adjustments for Automatic Lyrics Transcription
- Learning Recurrent Binary/Ternary Weights
- Robustness Testing of Language Understanding in Task-Oriented Dialog
- Continuous Transition: Improving Sample Efficiency for Continuous Control Problems via MixUp
- Speeding up Deep Model Training by Sharing Weights and Then Unsharing
- Greedy Optimization Provably Wins the Lottery: Logarithmic Number of Winning Tickets is Enough
- Distributed Deep Learning Strategies For Automatic Speech Recognition
- Guiding CTC Posterior Spike Timings for Improved Posterior Fusion and Knowledge Distillation
- Principled Hybrids of Generative and Discriminative Domain Adaptation
- Polyphone Disambiguation in Mandarin Chinese with Semi-Supervised Learning
- Transformer-based Online CTC/attention End-to-End Speech Recognition Architecture
- Stage-based Hyper-parameter Optimization for Deep Learning
- Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck
- A Dynamically Controlled Recurrent Neural Network for Modeling Dynamical Systems
- Label-efficient audio classification through multitask learning and self-supervision
- Cross-Modal Knowledge Distillation Method for Automatic Cued Speech Recognition
- MooseNet: A Trainable Metric for Synthesized Speech with a PLDA Module
- Speech Recognition With No Speech Or With Noisy Speech Beyond English
- Beyond Neural-on-Neural Approaches to Speaker Gender Protection
- Exploiting Large-scale Teacher-Student Training for On-device Acoustic Models
- Sequence-Level Knowledge Distillation for Model Compression of Attention-based Sequence-to-Sequence Speech Recognition
- Perceptual-based deep-learning denoiser as a defense against adversarial attacks on ASR systems
- CAAD 2018: Iterative Ensemble Adversarial Attack
- CNN-based MultiChannel End-to-End Speech Recognition for everyday home environments
- SimulLR: Simultaneous Lip Reading Transducer with Attention-Guided Adaptive Memory
- Cycle-consistency training for end-to-end speech recognition
- Back from the future: bidirectional CTC decoding using future information in speech recognition
- Unidirectional Memory-Self-Attention Transducer for Online Speech Recognition
- Extending Recurrent Neural Aligner for Streaming End-to-End Speech Recognition in Mandarin
- Faster, Simpler and More Accurate Hybrid ASR Systems Using Wordpieces
- Deep learning approaches for neural decoding: from CNNs to LSTMs and spikes to fMRI
- Hippo: Taming Hyper-parameter Optimization of Deep Learning with Stage Trees
- MASRI-HEADSET: A Maltese Corpus for Speech Recognition
- KoSpeech: Open-Source Toolkit for End-to-End Korean Speech Recognition
- Classifying the Cosmic-Ray Proton and Light Groups on the LHAASO-KM2A Experiment with the Graph Neural Network
- MIMO-SPEECH: End-to-End Multi-Channel Multi-Speaker Speech Recognition
- power-law nonlinearity with maximally uniform distribution criterion for improved neural network training in automatic speech recognition
- Attention-Based End-to-End Speech Recognition on Voice Search
- Deep Learning in Memristive Nanowire Networks
- KT-Speech-Crawler: Automatic Dataset Construction for Speech Recognition from YouTube Videos
- Latent Adversarial Defence with Boundary-guided Generation
- Variational hybridization and transformation for large inaccurate noisy-or networks
- Deep Triphone Embedding Improves Phoneme Recognition
- Quantized Adam with Error Feedback
- End-to-end Whispered Speech Recognition with Frequency-weighted Approaches and Pseudo Whisper Pre-training
- Multi-style Training for South African Call Centre Audio
- Accurate and Energy-Efficient Classification with Spiking Random Neural Network: Corrected and Expanded Version
- Character-Level Neural Translation for Multilingual Media Monitoring in the SUMMA Project
- Attention-based ASR with Lightweight and Dynamic Convolutions
- PyChain: A Fully Parallelized PyTorch Implementation of LF-MMI for End-to-End ASR
- Adaptive Clustering of Robust Semantic Representations for Adversarial Image Purification
- Phoneme-based Distribution Regularization for Speech Enhancement
- Two-stage Textual Knowledge Distillation for End-to-End Spoken Language Understanding
- Deep segmental phonetic posterior-grams based discovery of non-categories in L2 English speech
- SQWA: Stochastic Quantized Weight Averaging for Improving the Generalization Capability of Low-Precision Deep Neural Networks
- Improving Efficiency in Large-Scale Decentralized Distributed Training
- Explaining Adversarial Vulnerability with a Data Sparsity Hypothesis
- Label-Synchronous Speech-to-Text Alignment for ASR Using Forward and Backward Transformers
- AIBench: An Agile Domain-specific Benchmarking Methodology and an AI Benchmark Suite
- Enlarging Discriminative Power by Adding an Extra Class in Unsupervised Domain Adaptation
- Scale Calibrated Training: Improving Generalization of Deep Networks via Scale-Specific Normalization
- Streaming Transformer ASR with Blockwise Synchronous Beam Search
- Using multi-task learning to improve the performance of acoustic-to-word and conventional hybrid models
- Human-Machine Collaborative Design for Accelerated Design of Compact Deep Neural Networks for Autonomous Driving
- End-To-End Speech Recognition Using A High Rank LSTM-CTC Based Model
- Fully Dynamic Inference with Deep Neural Networks
- Learning Shared Encoding Representation for End-to-End Speech Recognition Models
- Halo: Learning Semantics-Aware Representations for Cross-Lingual Information Extraction
- From Semi-supervised to Almost-unsupervised Speech Recognition with Very-low Resource by Jointly Learning Phonetic Structures from Audio and Text Embeddings
- Harnessing Slow Dynamics in Neuromorphic Computation
- Asynchronous Decentralized Distributed Training of Acoustic Models
- Speech recognition for air traffic control via feature learning and end-to-end training
- Sparse Persistent RNNs: Squeezing Large Recurrent Networks On-Chip
- Time-Frequency Localization Using Deep Convolutional Maxout Neural Network in Persian Speech Recognition
- Curriculum Learning with Diversity for Supervised Computer Vision Tasks
- MaskCycleGAN-VC: Learning Non-parallel Voice Conversion with Filling in Frames
- Advancing Multi-Accented LSTM-CTC Speech Recognition using a Domain Specific Student-Teacher Learning Paradigm
- Adversarial Regression with Doubly Non-negative Weighting Matrices
- Multi-QuartzNet: Multi-Resolution Convolution for Speech Recognition with Multi-Layer Feature Fusion
- Towards thinner convolutional neural networks through Gradually Global Pruning
- The Marchex 2018 English Conversational Telephone Speech Recognition System
- Mixture of Expert/Imitator Networks: Scalable Semi-supervised Learning Framework
- Cascaded CNN-resBiLSTM-CTC: An End-to-End Acoustic Model For Speech Recognition
- Learning Neural Models for Natural Language Processing in the Face of Distributional Shift
- Online Evolutionary Batch Size Orchestration for Scheduling Deep Learning Workloads in GPU Clusters
- Formant Tracking Using Dilated Convolutional Networks Through Dense Connection with Gating Mechanism
- No-regret Non-convex Online Meta-Learning
- Transformer ASR with Contextual Block Processing
- CrowdSpeech and VoxDIY: Benchmark Datasets for Crowdsourced Audio Transcription
- Sampling-free Uncertainty Estimation in Gated Recurrent Units with Exponential Families
- Multi-Head Decoder for End-to-End Speech Recognition
- Experiments with Rich Regime Training for Deep Learning
- Forget the Learning Rate, Decay Loss
- Dynamic Spectrum Matching with One-shot Learning
- Manner of Articulation Detection using Connectionist Temporal Classification to Improve Automatic Speech Recognition Performance
- Beam Search Decoding using Manner of Articulation Detection Knowledge Derived from Connectionist Temporal Classification
- Dynamic Encoder Transducer: A Flexible Solution For Trading Off Accuracy For Latency
- Dual Head Adversarial Training
- Understanding Modern Techniques in Optimization: Frank-Wolfe, Nesterov's Momentum, and Polyak's Momentum
- A Sum-of-Ratios Multi-Dimensional-Knapsack Decomposition for DNN Resource Scheduling
- DTNN: Energy-efficient Inference with Dendrite Tree Inspired Neural Networks for Edge Vision Applications
- Automatic Documentation of ICD Codes with Far-Field Speech Recognition
- GaDei: On Scale-up Training As A Service For Deep Learning
- On Large Batch Training and Sharp Minima: A Fokker-Planck Perspective
- Convolutional Variational Autoencoders for Spectrogram Compression in Automatic Speech Recognition
- Joint Regularization on Activations and Weights for Efficient Neural Network Pruning
- Towards Learning to Speak and Hear Through Multi-Agent Communication over a Continuous Acoustic Channel
- End-to-End Bengali Speech Recognition
- LfEdNet: A Task-based Day-ahead Load Forecasting Model for Stochastic Economic Dispatch
- Improving OOV Detection and Resolution with External Language Models in Acoustic-to-Word ASR
- Semantic Communications for Speech Recognition
- Learning to Transfer Learn: Reinforcement Learning-Based Selection for Adaptive Transfer Learning
- A Novel Co-design Peta-scale Heterogeneous Cluster for Deep Learning Training
- Adaptive Recurrent Neural Network Based on Mixture Layer
- A Highly Efficient Distributed Deep Learning System For Automatic Speech Recognition
- Low Precision Floating-point Arithmetic for High Performance FPGA-based CNN Acceleration
- Latency-Controlled Neural Architecture Search for Streaming Speech Recognition
- Two-stage Training for Chinese Dialect Recognition
- Application of Word2vec in Phoneme Recognition
- Hard Class Rectification for Domain Adaptation
- Mapping the Internet: Modelling Entity Interactions in Complex Heterogeneous Networks
- Accelerating SGD for Distributed Deep-Learning Using Approximated Hessian Matrix
- Self-Adaptive Reconfigurable Arrays (SARA): Using ML to Assist Scaling GEMM Acceleration
- Distributed Parameter Estimation in Randomized One-hidden-layer Neural Networks
- Fast offline Transformer-based end-to-end automatic speech recognition for real-world applications
- Toward Streaming ASR with Non-Autoregressive Insertion-based Model
- Wav-BERT: Cooperative Acoustic and Linguistic Representation Learning for Low-Resource Speech Recognition
- Explaining the Attention Mechanism of End-to-End Speech Recognition Using Decision Trees
- Refining Automatic Speech Recognition System for older adults
- Unigram-Normalized Perplexity as a Language Model Performance Measure with Different Vocabulary Sizes
- A Brief Survey and an Application of Semantic Image Segmentation for Autonomous Driving
- Convergence analysis of neural networks for solving a free boundary system
- Leveraging Acoustic and Linguistic Embeddings from Pretrained speech and language Models for Intent Classification
- Improving speech recognition models with small samples for air traffic control systems
- Knowledge Distillation from BERT Transformer to Speech Transformer for Intent Classification