6 citations · 10 across the 6 of their papers we have counts for
6 papers · 1 filter
Multilingual and Fully Non-Autoregressive ASR with Large Language Model Fusion: A Comprehensive Study
W. Ronny Huang, Cyril Allauzen, Tongzhou Chen +7
In the era of large models, the autoregressive nature of decoding often results in latency serving as a significant bottleneck. We propose a non-autoregressive LM-fused ASR system…
Improving Joint Speech-Text Representations Without Alignment
Cal Peyser, Zhong Meng, Ke Hu +5
The last year has seen astonishing progress in text-prompted image generation premised on the idea of a cross-modal representation space in which the text and image domains are rep…
Mixture-of-Expert Conformer for Streaming Multilingual ASR
Ke Hu, Bo Li, Tara N. Sainath +2
End-to-end models with large capacity have significantly improved multilingual automatic speech recognition, but their computation cost poses challenges for on-device applications.…
Scaling Up Deliberation for Multilingual ASR
Ke Hu, Bo Li, Tara N. Sainath
Multilingual end-to-end automatic speech recognition models are attractive due to its simplicity in training and deployment. Recent work on large-scale training of such models has…
Streaming Align-Refine for Non-autoregressive Deliberation
Weiran Wang, Ke Hu, Tara N. Sainath
We propose a streaming non-autoregressive (non-AR) decoding algorithm to deliberate the hypothesis alignment of a streaming RNN-T model. Our algorithm facilitates a simple greedy d…
Adversarial Training for Multilingual Acoustic Modeling
Ke Hu, Hasim Sak, Hank Liao
Multilingual training has been shown to improve acoustic modeling performance by sharing and transferring knowledge in modeling different languages. Knowledge sharing is usually ac…