3 papers
cs.CL2026
Textual Entailment is not a Better Bias Metric than Token Probability
Virginia K. Felkner, Allison Lim, Jonathan May
Measurement of social bias in language models is typically by token probability (TP) metrics, which are broadly applicable but have been criticized for their distance from real-wor…
cs.CL2024
WinoQueer: A Community-in-the-Loop Benchmark for Anti-LGBTQ+ Bias in Large Language Models
Virginia K. Felkner, Ho-Chun Herbert Chang, Eugene Jang +1
We present WinoQueer: a benchmark specifically designed to measure whether large language models (LLMs) encode biases that are harmful to the LGBTQ+ community. The benchmark is com…
cs.CL2024
GPT is Not an Annotator: The Necessity of Human Annotation in Fairness Benchmark Construction
Virginia K. Felkner, Jennifer A. Thompson, Jonathan May
Social biases in LLMs are usually measured via bias benchmark datasets. Current benchmarks have limitations in scope, grounding, quality, and human effort required. Previous work h…