9 citations · 14 across the 4 of their papers we have counts for
4 papers
Mining Large-Scale Low-Resource Pronunciation Data From Wikipedia
Tania Chakraborty, Manasa Prasad, Theresa Breiner +2
Pronunciation modeling is a key task for building speech technology in new languages, and while solid grapheme-to-phoneme (G2P) mapping systems exist, language coverage can stand t…
Language ID in the Wild: Unexpected Challenges on the Path to a Thousand-Language Web Text Corpus
Isaac Caswell, Theresa Breiner, Daan van Esch +1
Large text corpora are increasingly important for a wide variety of Natural Language Processing (NLP) tasks, and automatic language identification (LangID) is a core technology nee…
Writing Across the World's Languages: Deep Internationalization for Gboard, the Google Keyboard
Daan van Esch, Elnaz Sarbar, Tamar Lucassen +6
This technical report describes our deep internationalization program for Gboard, the Google Keyboard. Today, Gboard supports 900+ language varieties across 70+ writing systems, an…
Automatic Keyboard Layout Design for Low-Resource Latin-Script Languages
Theresa Breiner, Chieu Nguyen, Daan van Esch +1
We present our approach to automatically designing and implementing keyboard layouts on mobile devices for typing low-resource languages written in the Latin script. For many speak…