113 citations · 118 across the 6 of their papers we have counts for
9 papers
Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages
Yu Zhang, Wei Han, James Qin +24
We introduce the Universal Speech Model (USM), a single large model that performs automatic speech recognition (ASR) across 100+ languages. This is achieved by pre-training the enc…
G-Augment: Searching for the Meta-Structure of Data Augmentation Policies for ASR
Gary Wang, Ekin D. Cubuk, Andrew Rosenberg +6
Data augmentation is a ubiquitous technique used to provide robustness to automatic speech recognition (ASR) training. However, even as so much of the ASR training process has beco…
Maestro-U: Leveraging joint speech-text representation learning for zero supervised speech ASR
Zhehuai Chen, Ankur Bapna, Andrew Rosenberg +4
Training state-of-the-art Automated Speech Recognition (ASR) models typically requires a substantial amount of transcribed speech. In this work, we demonstrate that a modality-matc…
Analysis of Self-Attention Head Diversity for Conformer-based Automatic Speech Recognition
Kartik Audhkhasi, Yinghui Huang, Bhuvana Ramabhadran +1
Attention layers are an integral part of modern end-to-end automatic speech recognition systems, for instance as part of the Transformer or Conformer architecture. Attention is typ…
Ask2Mask: Guided Data Selection for Masked Speech Modeling
Murali Karthick Baskar, Andrew Rosenberg, Bhuvana Ramabhadran +2
Masked speech modeling (MSM) methods such as wav2vec2 or w2v-BERT learn representations over speech frames which are randomly masked within an utterance. While these methods improv…
Injecting Text in Self-Supervised Speech Pretraining
Zhehuai Chen, Yu Zhang, Andrew Rosenberg +3
Self-supervised pretraining for Automated Speech Recognition (ASR) has shown varied degrees of success. In this paper, we propose to jointly learn representations during pretrainin…