activity
20172022
most citedDisembodied Machine Learning: On the Illusion of Objectivity in NLP

9 citations · 17 across the 6 of their papers we have counts for

collaborators

11 papers

cs.LG2022

A Federated Approach to Predicting Emojis in Hindi Tweets

Deep Gandhi, Jash Mehta, Nirali Parekh +3

The use of emojis affords a visual modality to, often private, textual communication. The task of predicting emojis however provides a challenge for machine learning as emoji use t…

cs.CL20221 cited

Back to the Future: On Potential Histories in NLP

Zeerak Talat, Anne Lauscher

Machine learning and NLP require the construction of datasets to train and fine-tune models. In this context, previous work has demonstrated the sensitivity of these data sets. For…

cs.CL20222 cited

Documenting Geographically and Contextually Diverse Data Sources: The BigScience Catalogue of Language Data and Resources

Angelina McMillan-Major, Zaid Alyafeai, Stella Biderman +15

In recent years, large-scale data collection efforts have prioritized the amount of data collected in order to improve the modeling capabilities of large language models. This prio…

cs.CL20211 cited

A Word on Machine Ethics: A Response to Jiang et al. (2021)

Zeerak Talat, Hagen Blix, Josef Valvoda +3

Ethics is one of the longest standing intellectual endeavors of humanity. In recent years, the fields of AI and NLP have attempted to wrangle with how learning systems that interac…

cs.CL2021

A Survey of Race, Racism, and Anti-Racism in NLP

Anjalie Field, Su Lin Blodgett, Zeerak Waseem +1

Despite inextricable ties between race and language, little work has considered race in NLP research and development. In this work, we survey 79 papers from the ACL anthology that…

cs.CL2021

Dynabench: Rethinking Benchmarking in NLP

Douwe Kiela, Max Bartolo, Yixin Nie +16

We introduce Dynabench, an open-source platform for dynamic dataset creation and model benchmarking. Dynabench runs in a web browser and supports human-and-model-in-the-loop datase…