papers

Publications (19)

cs.CL2021

Reducing Exposure Bias in Training Recurrent Neural Network Transducers

Xiaodong Cui, Brian Kingsbury, George Saon +2

When recurrent neural network transducers (RNNTs) are trained using the typical maximum likelihood criterion, the prediction network is trained only on ground truth label sequences…

eess.AS2021

Supervised and Unsupervised Approaches for Controlling Narrow Lexical Focus in Sequence-to-Sequence Speech Synthesis

Slava Shechtman, Raul Fernandez, David Haws

Although Sequence-to-Sequence (S2S) architectures have become state-of-the-art in speech synthesis, capable of generating outputs that approach the perceptual quality of natural sa…

math.ST2011

Semigroups and sequential importance sampling for multiway tables and beyond

Jing Xi, Shaoceng Wei, Feng Zhou +2

When an interval of integers between the lower bound l_i and the upper bounds u_i is the support of the marginal distribution n_i|(n_{i-1}, ...,n_1), Chen et al. 2005 noticed that…

eess.AS2025

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities

George Saon, Avihu Dekel, Alexander Brooks +21

Granite-speech LLMs are compact and efficient speech language models specifically designed for English ASR and automatic speech translation (AST). The models were trained by modali…

cs.SD2022

VQ-T: RNN Transducers using Vector-Quantized Prediction Network States

Jiatong Shi, George Saon, David Haws +2

Beam search, which is the dominant ASR decoding algorithm for end-to-end models, generates tree-structured hypotheses. However, recent studies have shown that decoding with hypothe…

math.ST2013

Markov degree of the three-state toric homogeneous Markov chain model

David Haws, Abraham Martín del Campo, Akimichi Takemura +1

We consider the three-state toric homogeneous Markov chain model (THMC) without loops and initial parameters. At time , the size of the design matrix is $6 \times 3\cdot 2^{T-1}…