Publications (19)
Reducing Exposure Bias in Training Recurrent Neural Network Transducers
Xiaodong Cui, Brian Kingsbury, George Saon +2
When recurrent neural network transducers (RNNTs) are trained using the typical maximum likelihood criterion, the prediction network is trained only on ground truth label sequences…
Supervised and Unsupervised Approaches for Controlling Narrow Lexical Focus in Sequence-to-Sequence Speech Synthesis
Slava Shechtman, Raul Fernandez, David Haws
Although Sequence-to-Sequence (S2S) architectures have become state-of-the-art in speech synthesis, capable of generating outputs that approach the perceptual quality of natural sa…
Semigroups and sequential importance sampling for multiway tables and beyond
Jing Xi, Shaoceng Wei, Feng Zhou +2
When an interval of integers between the lower bound l_i and the upper bounds u_i is the support of the marginal distribution n_i|(n_{i-1}, ...,n_1), Chen et al. 2005 noticed that…
Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities
George Saon, Avihu Dekel, Alexander Brooks +21
Granite-speech LLMs are compact and efficient speech language models specifically designed for English ASR and automatic speech translation (AST). The models were trained by modali…
VQ-T: RNN Transducers using Vector-Quantized Prediction Network States
Jiatong Shi, George Saon, David Haws +2
Beam search, which is the dominant ASR decoding algorithm for end-to-end models, generates tree-structured hypotheses. However, recent studies have shown that decoding with hypothe…
Markov degree of the three-state toric homogeneous Markov chain model
David Haws, Abraham MartÃn del Campo, Akimichi Takemura +1
We consider the three-state toric homogeneous Markov chain model (THMC) without loops and initial parameters. At time , the size of the design matrix is $6 \times 3\cdot 2^{T-1}…