activity
20182023
most citedGoogle USM: Scaling Automatic Speech Recognition Beyond 100 Languages

113 citations · 403 across the 28 of their papers we have counts for

collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL2023

Text Injection for Capitalization and Turn-Taking Prediction in Speech Models

Shaan Bijwadia, Shuo-yiin Chang, Weiran Wang +3

Text injection for automatic speech recognition (ASR), wherein unpaired text-only data is used to supplement paired audio-text data, has shown promising improvements for word error…

cs.CL2023

Improving Joint Speech-Text Representations Without Alignment

Cal Peyser, Zhong Meng, Ke Hu +5

The last year has seen astonishing progress in text-prompted image generation premised on the idea of a cross-modal representation space in which the text and image domains are rep…

cs.CL2023★ 113 cited

Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Yu Zhang, Wei Han, James Qin +24

We introduce the Universal Speech Model (USM), a single large model that performs automatic speech recognition (ASR) across 100+ languages. This is achieved by pre-training the enc…

cs.CL2021★ 1 cited

Factorized Neural Transducer for Efficient Language Model Adaptation

Xie Chen, Zhong Meng, Sarangarajan Parthasarathy +1

In recent years, end-to-end (E2E) based automatic speech recognition (ASR) systems have achieved great success due to their simplicity and promising performance. Neural Transducer…

cs.CL2020

On Minimum Word Error Rate Training of the Hybrid Autoregressive Transducer

Liang Lu, Zhong Meng, Naoyuki Kanda +2

Hybrid Autoregressive Transducer (HAT) is a recently proposed end-to-end acoustic model that extends the standard Recurrent Neural Network Transducer (RNN-T) for the purpose of the…

cs.CL2020

Serialized Output Training for End-to-End Overlapped Speech Recognition

Naoyuki Kanda, Yashesh Gaur, Xiaofei Wang +2

This paper proposes serialized output training (SOT), a novel framework for multi-speaker overlapped speech recognition based on an attention-based encoder-decoder approach. Instea…