collaborators

5 papers

cs.CL2025

Teaching and Critiquing Conceptualization and Operationalization in NLP

Vagrant Gautam

NLP researchers regularly invoke abstract concepts like "interpretability," "bias," "reasoning," and "stereotypes," without defining them. Each subfield has a shared understanding…

cs.CL2025

Colombian Waitresses y Jueces canadienses: Gender and Country Biases in Occupation Recommendations from LLMs

Elisa Forcada Rodríguez, Olatz Perez-de-Viñaspre, Jon Ander Campos +2

One of the goals of fairness research in NLP is to measure and mitigate stereotypical biases that are propagated by NLP systems. However, such work tends to focus on single axes of…

cs.CL2025

Agree to Disagree? A Meta-Evaluation of LLM Misgendering

Arjun Subramonian, Vagrant Gautam, Preethi Seshadri +3

Numerous methods have been proposed to measure LLM misgendering, including probability-based evaluations (e.g., automatically with templatic sentences) and generation-based evaluat…

cs.CL2025

Aligned Probing: Relating Toxic Behavior and Model Internals

Andreas Waldis, Vagrant Gautam, Anne Lauscher +2

We introduce aligned probing, a novel interpretability framework that aligns the behavior of language models (LMs), based on their outputs, and their internal representations (inte…

cs.CL2024

WinoPron: Revisiting English Winogender Schemas for Consistency, Coverage, and Grammatical Case

Vagrant Gautam, Julius Steuer, Eileen Bingert +3

While measuring bias and robustness in coreference resolution are important goals, such measurements are only as good as the tools we use to measure them. Winogender Schemas (Rudin…