activity
20152017
most citedDirect Acoustics-to-Word Models for English Conversational Speech Recognition

22 citations · 26 across the 2 of their papers we have counts for

collaborators

5 papers

cs.CL201722 cited

Direct Acoustics-to-Word Models for English Conversational Speech Recognition

Kartik Audhkhasi, Bhuvana Ramabhadran, George Saon +2

Recent work on end-to-end automatic speech recognition (ASR) has shown that the connectionist temporal classification (CTC) loss can be used to convert acoustics to phone or charac…

cs.CL20174 cited

English Conversational Telephone Speech Recognition by Humans and Machines

George Saon, Gakuto Kurata, Tom Sercu +9

One of the most difficult speech recognition tasks is accurate recognition of human to human communication. Advances in deep learning over the last few years have produced major sp…

cs.LG2016

Training variance and performance evaluation of neural networks in speech

Ewout van den Berg, Bhuvana Ramabhadran, Michael Picheny

In this work we study variance in the results of neural network training on a wide variety of configurations in automatic speech recognition. Although this variance itself is well…

cs.LG2016

A Comparison between Deep Neural Nets and Kernel Acoustic Models for Speech Recognition

Zhiyun Lu, Dong Guo, Alireza Bagheri Garakani +8

We study large-scale kernel methods for acoustic modeling and compare to DNNs on performance metrics related to both acoustic modeling and recognition. Measuring perplexity and fra…

cs.CL2015

The IBM 2015 English Conversational Telephone Speech Recognition System

George Saon, Hong-Kwang J. Kuo, Steven Rennie +1

We describe the latest improvements to the IBM English conversational telephone speech recognition system. Some of the techniques that were found beneficial are: maxout networks wi…