20 citations · 39 across the 6 of their papers we have counts for
7 papers · 1 filter
Processing South Asian Languages Written in the Latin Script: the Dakshina Dataset
Brian Roark, Lawrence Wolf-Sonkin, Christo Kirov +4
This paper describes the Dakshina dataset, a new resource consisting of text in both the Latin and native scripts for 12 South Asian languages. The dataset includes, for each langu…
On the Relationships Between the Grammatical Genders of Inanimate Nouns and Their Co-Occurring Adjectives and Verbs
Adina Williams, Ryan Cotterell, Lawrence Wolf-Sonkin +2
We use large-scale corpora in six different gendered languages, along with tools from NLP and information theory, to test whether there is a relationship between the grammatical ge…
Quantifying the Semantic Core of Gender Systems
Adina Williams, Ryan Cotterell, Lawrence Wolf-Sonkin +2
Many of the world's languages employ grammatical gender on the lexeme. For example, in Spanish, the word for 'house' (casa) is feminine, whereas the word for 'paper' (papel) is mas…
The SIGMORPHON 2019 Shared Task: Morphological Analysis in Context and Cross-Lingual Transfer for Inflection
Arya D. McCarthy, Ekaterina Vylomova, Shijie Wu +9
The SIGMORPHON 2019 shared task on cross-lingual transfer and contextual analysis in morphology examined transfer learning of inflection between 100 language pairs, as well as cont…
Unsupervised Discovery of Gendered Language through Latent-Variable Modeling
Alexander Hoyle, Wolf-Sonkin, Hanna Wallach +2
Studying the ways in which language is gendered has long been an area of interest in sociolinguistics. Studies have explored, for example, the speech of male and female characters…
Combining Sentiment Lexica with a Multi-View Variational Autoencoder
Alexander Hoyle, Lawrence Wolf-Sonkin, Hanna Wallach +2
When assigning quantitative labels to a dataset, different methodologies may rely on different scales. In particular, when assigning polarities to words in a sentiment lexicon, ann…