papers

Publications (11)

cs.CL2017

Cold Fusion: Training Seq2Seq Models Together with Language Models

Anuroop Sriram, Heewoo Jun, Sanjeev Satheesh +1

Sequence-to-sequence (Seq2Seq) models with attention have excelled at tasks which involve generating natural language sentences such as machine translation, image captioning and sp…

cs.CL2017

Deep Voice: Real-time Neural Text-to-Speech

Sercan O. Arik, Mike Chrzanowski, Adam Coates +9

We present Deep Voice, a production-quality text-to-speech system constructed entirely from deep neural networks. Deep Voice lays the groundwork for truly end-to-end neural speech…

cs.RO2015

An Empirical Evaluation of Deep Learning on Highway Driving

Brody Huval, Tao Wang, Sameep Tandon +10

Numerous groups have applied a variety of deep learning techniques to computer vision problems in highway perception scenarios. In this paper, we presented a number of empirical ev…

cs.CV2013

Deep learning for class-generic object detection

Brody Huval, Adam Coates, Andrew Ng

We investigate the use of deep neural networks for the novel task of class generic object detection. We show that neural networks originally designed for image recognition can be t…

cs.CL2017

Exploring Neural Transducers for End-to-End Speech Recognition

Eric Battenberg, Jitong Chen, Rewon Child +8

In this work, we perform an empirical comparison among the CTC, RNN-Transducer, and attention-based Seq2Seq models for end-to-end speech recognition. We show that, without any lang…

cs.CL2016

Active Learning for Speech Recognition: the Power of Gradients

Jiaji Huang, Rewon Child, Vinay Rao +3

In training speech recognition systems, labeling audio clips can be expensive, and not all data is equally valuable. Active learning aims to label only the most informative samples…

cs.CL2015

Deep Speech 2: End-to-End Speech Recognition in English and Mandarin

Dario Amodei, Rishita Anubhai, Eric Battenberg +31

We show that an end-to-end deep learning approach can be used to recognize either English or Mandarin Chinese speech--two vastly different languages. Because it replaces entire pip…

cs.CL2017

Reducing Bias in Production Speech Models

Eric Battenberg, Rewon Child, Adam Coates +13

Replacing hand-engineered pipelines with end-to-end deep learning systems has enabled strong results in applications like speech and object recognition. However, the causality and…

cs.CL2017

Convolutional Recurrent Neural Networks for Small-Footprint Keyword Spotting

Sercan O. Arik, Markus Kliegl, Rewon Child +5

Keyword spotting (KWS) constitutes a major component of human-technology interfaces. Maximizing the detection accuracy at a low false alarm (FA) rate, while minimizing the footprin…

cs.LG2017

Principled Hybrids of Generative and Discriminative Domain Adaptation

Han Zhao, Zhenyao Zhu, Junjie Hu +2

We propose a probabilistic framework for domain adaptation that blends both generative and discriminative modeling in a principled way. Under this framework, generative and discrim…

cs.CL2014

Deep Speech: Scaling up end-to-end speech recognition

Awni Hannun, Carl Case, Jared Casper +8

We present a state-of-the-art speech recognition system developed using end-to-end deep learning. Our architecture is significantly simpler than traditional speech systems, which r…