activity
20162024
most citedEnglish Broadcast News Speech Recognition by Humans and Machines

12 citations · 31 across the 26 of their papers we have counts for

collaborators

37 papers

cs.CL2024

Exploring the limits of decoder-only models trained on public speech recognition corpora

Ankit Gupta, George Saon, Brian Kingsbury

The emergence of industrial-scale speech recognition (ASR) models such as Whisper and USM, trained on 1M hours of weakly labelled and 12M hours of audio only proprietary data respe…

cs.CL2024

Joint Unsupervised and Supervised Training for Automatic Speech Recognition via Bilevel Optimization

A F M Saif, Xiaodong Cui, Han Shen +3

In this paper, we present a novel bilevel optimization-based training approach to training acoustic models for automatic speech recognition (ASR) tasks that we term {bi-level joint…

cs.CL2023

Comparison of Multilingual Self-Supervised and Weakly-Supervised Speech Pre-Training for Adaptation to Unseen Languages

Andrew Rouditchenko, Sameer Khurana, Samuel Thomas +6

Recent models such as XLS-R and Whisper have made multilingual speech technologies more accessible by pre-training on audio from around 100 spoken languages each. However, there ar…

cs.IT20231 cited

High-Dimensional Smoothed Entropy Estimation via Dimensionality Reduction

Kristjan Greenewald, Brian Kingsbury, Yuancheng Yu

We study the problem of overcoming exponential sample complexity in differential entropy estimation under Gaussian convolutions. Specifically, we consider the estimation of the dif…

cs.SD2022

VQ-T: RNN Transducers using Vector-Quantized Prediction Network States

Jiatong Shi, George Saon, David Haws +2

Beam search, which is the dominant ASR decoding algorithm for end-to-end models, generates tree-structured hypotheses. However, recent studies have shown that decoding with hypothe…

cs.CL2022

Towards End-to-End Integration of Dialog History for Improved Spoken Language Understanding

Vishal Sunder, Samuel Thomas, Hong-Kwang J. Kuo +3

Dialog history plays an important role in spoken language understanding (SLU) performance in a dialog system. For end-to-end (E2E) SLU, previous work has used dialog history in tex…