184 citations · 366 across the 8 of their papers we have counts for
19 papers
Textless Direct Speech-to-Speech Translation with Discrete Speech Representation
Xinjian Li, Ye Jia, Chung-Cheng Chiu
Research on speech-to-speech translation (S2ST) has progressed rapidly in recent years. Many end-to-end systems have been proposed and show advantages over conventional cascade sys…
XTREME-S: Evaluating Cross-lingual Speech Representations
Alexis Conneau, Ankur Bapna, Yu Zhang +16
We introduce XTREME-S, a new benchmark to evaluate universal cross-lingual speech representations in many languages. XTREME-S covers four task families: speech recognition, classif…
mSLAM: Massively multilingual joint pre-training for speech and text
Ankur Bapna, Colin Cherry, Yu Zhang +6
We present mSLAM, a multilingual Speech and LAnguage Model that learns cross-lingual cross-modal representations of speech and text by pre-training jointly on large amounts of unla…
SLAM: A Unified Encoder for Speech and Language Modeling via Speech-Text Joint Pre-Training
Ankur Bapna, Yu-an Chung, Nan Wu +7
Unsupervised pre-training is now the predominant approach for both text and speech understanding. Self-attention models pre-trained on large amounts of unannotated data have been h…
PnG BERT: Augmented BERT on Phonemes and Graphemes for Neural TTS
Ye Jia, Heiga Zen, Jonathan Shen +2
This paper introduces PnG BERT, a new encoder model for neural TTS. This model is augmented from the original BERT model, by taking both phoneme and grapheme representations of tex…
Parallel Tacotron: Non-Autoregressive and Controllable TTS
Isaac Elias, Heiga Zen, Jonathan Shen +4
Although neural end-to-end text-to-speech models can synthesize highly natural speech, there is still room for improvements to its efficiency and naturalness. This paper proposes a…