activity
20182022
most citedLingvo: a Modular and Scalable Framework for Sequence-to-Sequence Modeling

184 citations · 366 across the 8 of their papers we have counts for

collaborators
Showing cs.CLShow all

12 papers · 1 filter

cs.CL20221 cited

Textless Direct Speech-to-Speech Translation with Discrete Speech Representation

Xinjian Li, Ye Jia, Chung-Cheng Chiu

Research on speech-to-speech translation (S2ST) has progressed rapidly in recent years. Many end-to-end systems have been proposed and show advantages over conventional cascade sys…

cs.CL20221 cited

XTREME-S: Evaluating Cross-lingual Speech Representations

Alexis Conneau, Ankur Bapna, Yu Zhang +16

We introduce XTREME-S, a new benchmark to evaluate universal cross-lingual speech representations in many languages. XTREME-S covers four task families: speech recognition, classif…

cs.CL202259 cited

mSLAM: Massively multilingual joint pre-training for speech and text

Ankur Bapna, Colin Cherry, Yu Zhang +6

We present mSLAM, a multilingual Speech and LAnguage Model that learns cross-lingual cross-modal representations of speech and text by pre-training jointly on large amounts of unla…

cs.CL202150 cited

SLAM: A Unified Encoder for Speech and Language Modeling via Speech-Text Joint Pre-Training

Ankur Bapna, Yu-an Chung, Nan Wu +7

Unsupervised pre-training is now the predominant approach for both text and speech understanding. Self-attention models pre-trained on large amounts of unannotated data have been h…

cs.CL2021

PnG BERT: Augmented BERT on Phonemes and Graphemes for Neural TTS

Ye Jia, Heiga Zen, Jonathan Shen +2

This paper introduces PnG BERT, a new encoder model for neural TTS. This model is augmented from the original BERT model, by taking both phoneme and grapheme representations of tex…

cs.CL2019

Speech Recognition with Augmented Synthesized Speech

Andrew Rosenberg, Yu Zhang, Bhuvana Ramabhadran +4

Recent success of the Tacotron speech synthesis architecture and its variants in producing natural sounding multi-speaker synthesized speech has raised the exciting possibility of…