4 papers
Machines Do See Color: A Guideline to Classify Different Forms of Racist Discourse in Large Corpora
Diana Davila Gordillo, Joan C. Timoneda, Sebastian Vallejo Vera
Current methods to identify and classify racist language in text rely on small-n qualitative approaches or large-n approaches focusing exclusively on overt forms of racist discours…
The Synthetic Imputation Approach: Generating Optimal Synthetic Texts For Underrepresented Categories In Supervised Classification Tasks
Joan C. Timoneda
Encoder-decoder Large Language Models (LLMs), such as BERT and RoBERTa, require that all categories in an annotation task be sufficiently represented in the training data for optim…
Memory Is All You Need: Testing How Model Memory Affects LLM Performance in Annotation Tasks
Joan C. Timoneda, Sebastián Vallejo Vera
Generative Large Language Models (LLMs) have shown promising results in text annotation using zero-shot and few-shot learning. Yet these approaches do not allow the model to retain…
Identifying the sources of ideological bias in GPT models through linguistic variation in output
Christina Walker, Joan C. Timoneda
Extant work shows that generative AI models such as GPT-3.5 and 4 perpetuate social stereotypes and biases. One concerning but less explored source of bias is ideology. Do GPT mode…