3 papers
cs.CL2023
Audio-AdapterFusion: A Task-ID-free Approach for Efficient and Non-Destructive Multi-task Speech Recognition
Hillary Ngai, Rohan Agrawal, Neeraj Gaur +3
Adapters are an efficient, composable alternative to full fine-tuning of pre-trained models and help scale the deployment of large ASR models to many tasks. In practice, a task ID…
cs.CL2023
A Comparison of Semi-Supervised Learning Techniques for Streaming ASR at Scale
Cal Peyser, Michael Picheny, Kyunghyun Cho +3
Unpaired text and audio injection have emerged as dominant methods for improving ASR performance in the absence of a large labeled corpus. However, little guidance exists on deploy…
cs.CL2023
Dual Learning for Large Vocabulary On-Device ASR
Cal Peyser, Ronny Huang, Tara Sainath +3
Dual learning is a paradigm for semi-supervised machine learning that seeks to leverage unsupervised data by solving two opposite tasks at once. In this scheme, each model is used…