12 citations · 31 across the 26 of their papers we have counts for
37 papers
Exploring the limits of decoder-only models trained on public speech recognition corpora
Ankit Gupta, George Saon, Brian Kingsbury
The emergence of industrial-scale speech recognition (ASR) models such as Whisper and USM, trained on 1M hours of weakly labelled and 12M hours of audio only proprietary data respe…
Joint Unsupervised and Supervised Training for Automatic Speech Recognition via Bilevel Optimization
A F M Saif, Xiaodong Cui, Han Shen +3
In this paper, we present a novel bilevel optimization-based training approach to training acoustic models for automatic speech recognition (ASR) tasks that we term {bi-level joint…
Comparison of Multilingual Self-Supervised and Weakly-Supervised Speech Pre-Training for Adaptation to Unseen Languages
Andrew Rouditchenko, Sameer Khurana, Samuel Thomas +6
Recent models such as XLS-R and Whisper have made multilingual speech technologies more accessible by pre-training on audio from around 100 spoken languages each. However, there ar…
High-Dimensional Smoothed Entropy Estimation via Dimensionality Reduction
Kristjan Greenewald, Brian Kingsbury, Yuancheng Yu
We study the problem of overcoming exponential sample complexity in differential entropy estimation under Gaussian convolutions. Specifically, we consider the estimation of the dif…
VQ-T: RNN Transducers using Vector-Quantized Prediction Network States
Jiatong Shi, George Saon, David Haws +2
Beam search, which is the dominant ASR decoding algorithm for end-to-end models, generates tree-structured hypotheses. However, recent studies have shown that decoding with hypothe…
Towards End-to-End Integration of Dialog History for Improved Spoken Language Understanding
Vishal Sunder, Samuel Thomas, Hong-Kwang J. Kuo +3
Dialog history plays an important role in spoken language understanding (SLU) performance in a dialog system. For end-to-end (E2E) SLU, previous work has used dialog history in tex…