293 citations · 946 across the 20 of their papers we have counts for
23 papers · 1 filter
Maestro-U: Leveraging joint speech-text representation learning for zero supervised speech ASR
Zhehuai Chen, Ankur Bapna, Andrew Rosenberg +4
Training state-of-the-art Automated Speech Recognition (ASR) models typically requires a substantial amount of transcribed speech. In this work, we demonstrate that a modality-matc…
JOIST: A Joint Speech and Text Streaming Model For ASR
Tara N. Sainath, Rohit Prabhavalkar, Ankur Bapna +6
We present JOIST, an algorithm to train a streaming, cascaded, encoder end-to-end (E2E) model with both speech-text paired inputs, and text-only unpaired inputs. Unlike previous wo…
FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech
Alexis Conneau, Min Ma, Simran Khanuja +6
We introduce FLEURS, the Few-shot Learning Evaluation of Universal Representations of Speech benchmark. FLEURS is an n-way parallel speech dataset in 102 languages built on top of…
XTREME-S: Evaluating Cross-lingual Speech Representations
Alexis Conneau, Ankur Bapna, Yu Zhang +16
We introduce XTREME-S, a new benchmark to evaluate universal cross-lingual speech representations in many languages. XTREME-S covers four task families: speech recognition, classif…
Multilingual Mix: Example Interpolation Improves Multilingual Neural Machine Translation
Yong Cheng, Ankur Bapna, Orhan Firat +3
Multilingual neural machine translation models are trained to maximize the likelihood of a mix of examples drawn from multiple language pairs. The dominant inductive bias applied t…
Examining Scaling and Transfer of Language Model Architectures for Machine Translation
Biao Zhang, Behrooz Ghorbani, Ankur Bapna +4
Natural language understanding and generation models follow one of the two dominant architectural paradigms: language models (LMs) that process concatenated sequences in a single s…