20 citations · 20 across the 1 of their papers we have counts for
2 papers
cs.CL2020
Google Crowdsourced Speech Corpora and Related Open-Source Resources for Low-Resource Languages and Dialects: An Overview
Alena Butryna, Shan-Hui Cathy Chu, Isin Demirsahin +18
This paper presents an overview of a program designed to address the growing need for developing freely available speech resources for under-represented languages. At present we ha…
cs.CL2020★ 20 cited
Processing South Asian Languages Written in the Latin Script: the Dakshina Dataset
Brian Roark, Lawrence Wolf-Sonkin, Christo Kirov +4
This paper describes the Dakshina dataset, a new resource consisting of text in both the Latin and native scripts for 12 South Asian languages. The dataset includes, for each langu…