activity
20152023
most citedDirect Acoustics-to-Word Models for English Conversational Speech Recognition

22 citations · 51 across the 20 of their papers we have counts for

collaborators
Showing cs.CLShow all

20 papers · 1 filter

cs.CL2022

Improving Generalization of Deep Neural Network Acoustic Models with Length Perturbation and N-best Based Label Smoothing

Xiaodong Cui, George Saon, Tohru Nagano +4

We introduce two techniques, length perturbation and n-best based label smoothing, to improve generalization of deep neural network (DNN) acoustic models for automatic speech recog…

cs.CL2022

Towards Reducing the Need for Speech Training Data To Build Spoken Language Understanding Systems

Samuel Thomas, Hong-Kwang J. Kuo, Brian Kingsbury +1

The lack of speech data annotated with labels required for spoken language understanding (SLU) is often a major hurdle in building end-to-end (E2E) systems that can directly proces…

cs.CL2022

Integrating Text Inputs For Training and Adapting RNN Transducer ASR Models

Samuel Thomas, Brian Kingsbury, George Saon +1

Compared to hybrid automatic speech recognition (ASR) systems that use a modular architecture in which each component can be independently adapted to a new domain, recent end-to-en…

cs.CL2022

Improving End-to-End Models for Set Prediction in Spoken Language Understanding

Hong-Kwang J. Kuo, Zoltan Tuske, Samuel Thomas +2

The goal of spoken language understanding (SLU) systems is to determine the meaning of the input speech signal, unlike speech recognition which aims to produce verbatim transcripts…

cs.CL20211 cited

Asynchronous Decentralized Distributed Training of Acoustic Models

Xiaodong Cui, Wei Zhang, Abdullah Kayi +5

Large-scale distributed training of deep acoustic models plays an important role in today's high-performance automatic speech recognition (ASR). In this paper we investigate a vari…

cs.CL2021

4-bit Quantization of LSTM-based Speech Recognition Models

Andrea Fasoli, Chia-Yu Chen, Mauricio Serrano +9

We investigate the impact of aggressive low-precision representations of weights and activations in two families of large LSTM-based architectures for Automatic Speech Recognition…