297 citations · 1.5k across the 37 of their papers we have counts for
14 papers · 1 filter
MuSLAM: Multitask, Multilingual Speech and Language Models
Yong Cheng, Yu Zhang, Melvin Johnson +2
We present MuSLAM, a multilingual sequence-to-sequence model pre-trained jointly on unlabeled speech, unlabeled text and supervised data spanning Automatic Speech Recognition…
Maestro-U: Leveraging joint speech-text representation learning for zero supervised speech ASR
Zhehuai Chen, Ankur Bapna, Andrew Rosenberg +4
Training state-of-the-art Automated Speech Recognition (ASR) models typically requires a substantial amount of transcribed speech. In this work, we demonstrate that a modality-matc…
JOIST: A Joint Speech and Text Streaming Model For ASR
Tara N. Sainath, Rohit Prabhavalkar, Ankur Bapna +6
We present JOIST, an algorithm to train a streaming, cascaded, encoder end-to-end (E2E) model with both speech-text paired inputs, and text-only unpaired inputs. Unlike previous wo…
Virtuoso: Massive Multilingual Speech-Text Joint Semi-Supervised Learning for Text-To-Speech
Takaaki Saeki, Heiga Zen, Zhehuai Chen +6
This paper proposes Virtuoso, a massively multilingual speech-text joint semi-supervised learning framework for text-to-speech synthesis (TTS) models. Existing multilingual TTS typ…
SQuId: Measuring Speech Naturalness in Many Languages
Thibault Sellam, Ankur Bapna, Joshua Camp +3
Much of text-to-speech research relies on human evaluation, which incurs heavy costs and slows down the development process. The problem is particularly acute in heavily multilingu…
FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech
Alexis Conneau, Min Ma, Simran Khanuja +6
We introduce FLEURS, the Few-shot Learning Evaluation of Universal Representations of Speech benchmark. FLEURS is an n-way parallel speech dataset in 102 languages built on top of…