Publications (11)
Cold Fusion: Training Seq2Seq Models Together with Language Models
Anuroop Sriram, Heewoo Jun, Sanjeev Satheesh +1
Sequence-to-sequence (Seq2Seq) models with attention have excelled at tasks which involve generating natural language sentences such as machine translation, image captioning and sp…
Deep Voice: Real-time Neural Text-to-Speech
Sercan O. Arik, Mike Chrzanowski, Adam Coates +9
We present Deep Voice, a production-quality text-to-speech system constructed entirely from deep neural networks. Deep Voice lays the groundwork for truly end-to-end neural speech…
An Empirical Evaluation of Deep Learning on Highway Driving
Brody Huval, Tao Wang, Sameep Tandon +10
Numerous groups have applied a variety of deep learning techniques to computer vision problems in highway perception scenarios. In this paper, we presented a number of empirical ev…
Deep learning for class-generic object detection
Brody Huval, Adam Coates, Andrew Ng
We investigate the use of deep neural networks for the novel task of class generic object detection. We show that neural networks originally designed for image recognition can be t…
Exploring Neural Transducers for End-to-End Speech Recognition
Eric Battenberg, Jitong Chen, Rewon Child +8
In this work, we perform an empirical comparison among the CTC, RNN-Transducer, and attention-based Seq2Seq models for end-to-end speech recognition. We show that, without any lang…
Active Learning for Speech Recognition: the Power of Gradients
Jiaji Huang, Rewon Child, Vinay Rao +3
In training speech recognition systems, labeling audio clips can be expensive, and not all data is equally valuable. Active learning aims to label only the most informative samples…
Deep Speech 2: End-to-End Speech Recognition in English and Mandarin
Dario Amodei, Rishita Anubhai, Eric Battenberg +31
We show that an end-to-end deep learning approach can be used to recognize either English or Mandarin Chinese speech--two vastly different languages. Because it replaces entire pip…
Reducing Bias in Production Speech Models
Eric Battenberg, Rewon Child, Adam Coates +13
Replacing hand-engineered pipelines with end-to-end deep learning systems has enabled strong results in applications like speech and object recognition. However, the causality and…
Convolutional Recurrent Neural Networks for Small-Footprint Keyword Spotting
Sercan O. Arik, Markus Kliegl, Rewon Child +5
Keyword spotting (KWS) constitutes a major component of human-technology interfaces. Maximizing the detection accuracy at a low false alarm (FA) rate, while minimizing the footprin…
Principled Hybrids of Generative and Discriminative Domain Adaptation
Han Zhao, Zhenyao Zhu, Junjie Hu +2
We propose a probabilistic framework for domain adaptation that blends both generative and discriminative modeling in a principled way. Under this framework, generative and discrim…
Deep Speech: Scaling up end-to-end speech recognition
Awni Hannun, Carl Case, Jared Casper +8
We present a state-of-the-art speech recognition system developed using end-to-end deep learning. Our architecture is significantly simpler than traditional speech systems, which r…