293 citations · 790 across the 11 of their papers we have counts for
19 papers · 1 filter
XTREME-S: Evaluating Cross-lingual Speech Representations
Alexis Conneau, Ankur Bapna, Yu Zhang +16
We introduce XTREME-S, a new benchmark to evaluate universal cross-lingual speech representations in many languages. XTREME-S covers four task families: speech recognition, classif…
mSLAM: Massively multilingual joint pre-training for speech and text
Ankur Bapna, Colin Cherry, Yu Zhang +6
We present mSLAM, a multilingual Speech and LAnguage Model that learns cross-lingual cross-modal representations of speech and text by pre-training jointly on large amounts of unla…
SLAM: A Unified Encoder for Speech and Language Modeling via Speech-Text Joint Pre-Training
Ankur Bapna, Yu-an Chung, Nan Wu +7
Unsupervised pre-training is now the predominant approach for both text and speech understanding. Self-attention models pre-trained on large amounts of unannotated data have been h…
HintedBT: Augmenting Back-Translation with Quality and Transliteration Hints
Sahana Ramnath, Melvin Johnson, Abhirut Gupta +1
Back-translation (BT) of target monolingual corpora is a widely used data augmentation strategy for neural machine translation (NMT), especially for low-resource language pairs. To…
XTREME-R: Towards More Challenging and Nuanced Multilingual Evaluation
Sebastian Ruder, Noah Constant, Jan Botha +8
Machine learning has brought striking advances in multilingual natural language processing capabilities over the past year. For example, the latest techniques have improved the sta…
They, Them, Theirs: Rewriting with Gender-Neutral English
Tony Sun, Kellie Webster, Apu Shah +2
Responsible development of technology involves applications being inclusive of the diverse set of users they hope to support. An important part of this is understanding the many wa…