20 citations · 34 across the 4 of their papers we have counts for
15 papers
SIGTYP 2021 Shared Task: Robust Spoken Language Identification
Elizabeth Salesky, Badr M. Abdullah, Sabrina J. Mielke +6
While language identification is a fundamental speech and language processing task, for many languages and language families it remains a challenging task. For many low-resource an…
SIGTYP 2020 Shared Task: Prediction of Typological Features
Johannes Bjerva, Elizabeth Salesky, Sabrina J. Mielke +6
Typological knowledge bases (KBs) such as WALS (Dryer and Haspelmath, 2013) contain information about linguistic properties of the world's languages. They have been shown to be use…
SIGMORPHON 2020 Shared Task 0: Typologically Diverse Morphological Inflection
Ekaterina Vylomova, Jennifer White, Elizabeth Salesky +25
A broad goal in natural language processing (NLP) is to develop a system that has the capacity to process any natural language. Most systems, however, are developed using data from…
Processing South Asian Languages Written in the Latin Script: the Dakshina Dataset
Brian Roark, Lawrence Wolf-Sonkin, Christo Kirov +4
This paper describes the Dakshina dataset, a new resource consisting of text in both the Latin and native scripts for 12 South Asian languages. The dataset includes, for each langu…
It's Easier to Translate out of English than into it: Measuring Neural Translation Difficulty by Cross-Mutual Information
Emanuele Bugliarello, Sabrina J. Mielke, Antonios Anastasopoulos +2
The performance of neural machine translation systems is commonly evaluated in terms of BLEU. However, due to its reliance on target language properties and generation, the BLEU me…
Tired of Topic Models? Clusters of Pretrained Word Embeddings Make for Fast and Good Topics too!
Suzanna Sia, Ayush Dalmia, Sabrina J. Mielke
Topic models are a useful analysis tool to uncover the underlying themes within document collections. The dominant approach is to use probabilistic topic models that posit a genera…