activity
20152017
most citedDirect Acoustics-to-Word Models for English Conversational Speech Recognition

22 citations · 33 across the 5 of their papers we have counts for

collaborators

7 papers

cs.CL2017

Building competitive direct acoustics-to-word models for English conversational speech recognition

Kartik Audhkhasi, Brian Kingsbury, Bhuvana Ramabhadran +2

Direct acoustics-to-word (A2W) models in the end-to-end paradigm have received increasing attention compared to conventional sub-word based automatic speech recognition models usin…

cs.CL20176 cited

Embedding-Based Speaker Adaptive Training of Deep Neural Networks

Xiaodong Cui, Vaibhava Goel, George Saon

An embedding-based speaker adaptive training (SAT) approach is proposed and investigated in this paper for deep neural network acoustic modeling. In this approach, speaker embeddin…

cs.CL20171 cited

Language Modeling with Highway LSTM

Gakuto Kurata, Bhuvana Ramabhadran, George Saon +1

Language models (LMs) based on Long Short Term Memory (LSTM) have shown good gains in many automatic speech recognition tasks. In this paper, we extend an LSTM by adding highway ne…

cs.CL201722 cited

Direct Acoustics-to-Word Models for English Conversational Speech Recognition

Kartik Audhkhasi, Bhuvana Ramabhadran, George Saon +2

Recent work on end-to-end automatic speech recognition (ASR) has shown that the connectionist temporal classification (CTC) loss can be used to convert acoustics to phone or charac…

cs.CL20174 cited

English Conversational Telephone Speech Recognition by Humans and Machines

George Saon, Gakuto Kurata, Tom Sercu +9

One of the most difficult speech recognition tasks is accurate recognition of human to human communication. Advances in deep learning over the last few years have produced major sp…

cs.CL2016

The IBM 2016 English Conversational Telephone Speech Recognition System

George Saon, Tom Sercu, Steven Rennie +1

We describe a collection of acoustic and language modeling techniques that lowered the word error rate of our English conversational telephone LVCSR system to a record 6.6% on the…