activity
20182023
most citedGoogle USM: Scaling Automatic Speech Recognition Beyond 100 Languages

113 citations · 118 across the 6 of their papers we have counts for

collaborators

9 papers

cs.CL2023★ 113 cited

Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Yu Zhang, Wei Han, James Qin +24

We introduce the Universal Speech Model (USM), a single large model that performs automatic speech recognition (ASR) across 100+ languages. This is achieved by pre-training the enc…

cs.LG2022

G-Augment: Searching for the Meta-Structure of Data Augmentation Policies for ASR

Gary Wang, Ekin D. Cubuk, Andrew Rosenberg +6

Data augmentation is a ubiquitous technique used to provide robustness to automatic speech recognition (ASR) training. However, even as so much of the ASR training process has beco…

cs.CL2022

Maestro-U: Leveraging joint speech-text representation learning for zero supervised speech ASR

Zhehuai Chen, Ankur Bapna, Andrew Rosenberg +4

Training state-of-the-art Automated Speech Recognition (ASR) models typically requires a substantial amount of transcribed speech. In this work, we demonstrate that a modality-matc…

cs.CL2022

Analysis of Self-Attention Head Diversity for Conformer-based Automatic Speech Recognition

Kartik Audhkhasi, Yinghui Huang, Bhuvana Ramabhadran +1

Attention layers are an integral part of modern end-to-end automatic speech recognition systems, for instance as part of the Transformer or Conformer architecture. Attention is typ…

cs.SD2022★ 1 cited

Ask2Mask: Guided Data Selection for Masked Speech Modeling

Murali Karthick Baskar, Andrew Rosenberg, Bhuvana Ramabhadran +2

Masked speech modeling (MSM) methods such as wav2vec2 or w2v-BERT learn representations over speech frames which are randomly masked within an utterance. While these methods improv…

cs.CL2021★ 4 cited

Injecting Text in Self-Supervised Speech Pretraining

Zhehuai Chen, Yu Zhang, Andrew Rosenberg +3

Self-supervised pretraining for Automated Speech Recognition (ASR) has shown varied degrees of success. In this paper, we propose to jointly learn representations during pretrainin…