activity
20202022
most citedAccented Speech Recognition: Benchmarking, Pre-training, and Diverse Data

12 citations · 13 across the 5 of their papers we have counts for

collaborators

5 papers

eess.AS202212 cited

Accented Speech Recognition: Benchmarking, Pre-training, and Diverse Data

Alëna Aksënova, Zhehuai Chen, Chung-Cheng Chiu +8

Building inclusive speech recognition systems is a crucial step towards developing technologies that speakers of all language varieties can use. Therefore, ASR systems must work fo…

cs.CL20221 cited

XTREME-S: Evaluating Cross-lingual Speech Representations

Alexis Conneau, Ankur Bapna, Yu Zhang +16

We introduce XTREME-S, a new benchmark to evaluate universal cross-lingual speech representations in many languages. XTREME-S covers four task families: speech recognition, classif…

cs.CL2022

Handling Compounding in Mobile Keyboard Input

Andreas Kabel, Keith Hall, Tom Ouyang +3

This paper proposes a framework to improve the typing experience of mobile users in morphologically rich languages. Smartphone keyboards typically support features such as input de…

cs.CL2021

Mining Large-Scale Low-Resource Pronunciation Data From Wikipedia

Tania Chakraborty, Manasa Prasad, Theresa Breiner +2

Pronunciation modeling is a key task for building speech technology in new languages, and while solid grapheme-to-phoneme (G2P) mapping systems exist, language coverage can stand t…

cs.CL2020

Language ID in the Wild: Unexpected Challenges on the Path to a Thousand-Language Web Text Corpus

Isaac Caswell, Theresa Breiner, Daan van Esch +1

Large text corpora are increasingly important for a wide variety of Natural Language Processing (NLP) tasks, and automatic language identification (LangID) is a core technology nee…