5 papers
Text Injection for Neural Contextual Biasing
Zhong Meng, Zelin Wu, Rohit Prabhavalkar +5
Neural contextual biasing effectively improves automatic speech recognition (ASR) for crucial phrases within a speaker's context, particularly those that are infrequent in the trai…
Improving Joint Speech-Text Representations Without Alignment
Cal Peyser, Zhong Meng, Ke Hu +5
The last year has seen astonishing progress in text-prompted image generation premised on the idea of a cross-modal representation space in which the text and image domains are rep…
A Comparison of Semi-Supervised Learning Techniques for Streaming ASR at Scale
Cal Peyser, Michael Picheny, Kyunghyun Cho +3
Unpaired text and audio injection have emerged as dominant methods for improving ASR performance in the absence of a large labeled corpus. However, little guidance exists on deploy…
Dual Learning for Large Vocabulary On-Device ASR
Cal Peyser, Ronny Huang, Tara Sainath +3
Dual learning is a paradigm for semi-supervised machine learning that seeks to leverage unsupervised data by solving two opposite tasks at once. In this scheme, each model is used…
Towards Disentangled Speech Representations
Cal Peyser, Ronny Huang Andrew Rosenberg Tara N. Sainath, Michael Picheny +1
The careful construction of audio representations has become a dominant feature in the design of approaches to many speech tasks. Increasingly, such approaches have emphasized "dis…