4 papers
Standard Language Ideology in AI-Generated Language
Genevieve Smith, Eve Fleisig, Ishita Rustagi +1
Large language models (LLMs) generate text that reinforces standard language ideology: a bias towards certain language varieties that are granted more prestige, authority, and legi…
Characterizing Language Use in a Collaborative Situated Game
Nicholas Tomlin, Naitian Zhou, Eve Fleisig +10
Cooperative video games, where multiple participants must coordinate by communicating and reasoning under uncertainty in complex environments, yield a rich source of language data.…
Balancing Quality and Variation: Spam Filtering Distorts Data Label Distributions
Eve Fleisig, Matthias Orlikowski, Philipp Cimiano +1
For machine learning datasets to accurately represent diverse opinions in a population, they must preserve variation in data labels while filtering out spam or low-quality response…
Accurate and Data-Efficient Toxicity Prediction when Annotators Disagree
Harbani Jaggi, Kashyap Murali, Eve Fleisig +1
When annotators disagree, predicting the labels given by individual annotators can capture nuances overlooked by traditional label aggregation. We introduce three approaches to pre…