387 citations · 756 across the 29 of their papers we have counts for
52 papers
Efficient Adapters for Giant Speech Models
Nanxin Chen, Izhak Shafran, Yu Zhang +4
Large pre-trained speech models are widely used as the de-facto paradigm, especially in scenarios when there is a limited amount of labeled data available. However, finetuning all…
Accelerating RNN-T Training and Inference Using CTC guidance
Yongqiang Wang, Zhehuai Chen, Chengjian Zheng +3
We propose a novel method to accelerate training and inference process of recurrent neural network transducer (RNN-T) based on the guidance from a co-trained connectionist temporal…
Residual Adapters for Few-Shot Text-to-Speech Speaker Adaptation
Nobuyuki Morioka, Heiga Zen, Nanxin Chen +2
Adapting a neural text-to-speech (TTS) model to a target speaker typically involves fine-tuning most if not all of the parameters of a pretrained multi-speaker backbone model. Howe…
Comparison of Soft and Hard Target RNN-T Distillation for Large-scale ASR
Dongseong Hwang, Khe Chai Sim, Yu Zhang +1
Knowledge distillation is an effective machine learning technique to transfer knowledge from a teacher model to a smaller student model, especially with unlabeled data. In this pap…
Maestro-U: Leveraging joint speech-text representation learning for zero supervised speech ASR
Zhehuai Chen, Ankur Bapna, Andrew Rosenberg +4
Training state-of-the-art Automated Speech Recognition (ASR) models typically requires a substantial amount of transcribed speech. In this work, we demonstrate that a modality-matc…
FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech
Alexis Conneau, Min Ma, Simran Khanuja +6
We introduce FLEURS, the Few-shot Learning Evaluation of Universal Representations of Speech benchmark. FLEURS is an n-way parallel speech dataset in 102 languages built on top of…