activity
20172025
most citedGemini 1.5: Unlocking multimodal understanding across millions of tokens of context

297 citations · 1.5k across the 37 of their papers we have counts for

collaborators
Showing 2022Show all

14 papers · 1 filter

cs.CL2022★ 5 cited

MuSLAM: Multitask, Multilingual Speech and Language Models

Yong Cheng, Yu Zhang, Melvin Johnson +2

We present MuSLAM, a multilingual sequence-to-sequence model pre-trained jointly on unlabeled speech, unlabeled text and supervised data spanning Automatic Speech Recognition…

cs.CL2022

Maestro-U: Leveraging joint speech-text representation learning for zero supervised speech ASR

Zhehuai Chen, Ankur Bapna, Andrew Rosenberg +4

Training state-of-the-art Automated Speech Recognition (ASR) models typically requires a substantial amount of transcribed speech. In this work, we demonstrate that a modality-matc…

cs.CL2022★ 1 cited

JOIST: A Joint Speech and Text Streaming Model For ASR

Tara N. Sainath, Rohit Prabhavalkar, Ankur Bapna +6

We present JOIST, an algorithm to train a streaming, cascaded, encoder end-to-end (E2E) model with both speech-text paired inputs, and text-only unpaired inputs. Unlike previous wo…

cs.SD2022

Virtuoso: Massive Multilingual Speech-Text Joint Semi-Supervised Learning for Text-To-Speech

Takaaki Saeki, Heiga Zen, Zhehuai Chen +6

This paper proposes Virtuoso, a massively multilingual speech-text joint semi-supervised learning framework for text-to-speech synthesis (TTS) models. Existing multilingual TTS typ…

cs.CL2022★ 17 cited

SQuId: Measuring Speech Naturalness in Many Languages

Thibault Sellam, Ankur Bapna, Joshua Camp +3

Much of text-to-speech research relies on human evaluation, which incurs heavy costs and slows down the development process. The problem is particularly acute in heavily multilingu…

cs.CL2022★ 15 cited

FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech

Alexis Conneau, Min Ma, Simran Khanuja +6

We introduce FLEURS, the Few-shot Learning Evaluation of Universal Representations of Speech benchmark. FLEURS is an n-way parallel speech dataset in 102 languages built on top of…