6 papers
Evaluating Implicit Biases in LLM Reasoning through Logic Grid Puzzles
Fatima Jahara, Mark Dredze, Sharon Levy
While recent safety guardrails effectively suppress overtly biased outputs, subtler forms of social bias emerge during complex logical reasoning tasks that evade current evaluation…
Knowing But Not Doing: Convergent Morality and Divergent Action in LLMs
Jen-tse Huang, Jiantong Qin, Xueli Qiu +3
Value alignment is central to the development of safe and socially compatible artificial intelligence. However, how Large Language Models (LLMs) represent and enact human values in…
Characterizing Selective Refusal Bias in Large Language Models
Adel Khorramrouz, Sharon Levy
Safety guardrails in large language models(LLMs) are developed to prevent malicious users from generating toxic content at a large scale. However, these measures can inadvertently…
Making FETCH! Happen: Finding Emergent Dog Whistles Through Common Habitats
Kuleen Sasse, Carlos Aguirre, Isabel Cachola +2
WARNING: This paper contains content that maybe upsetting or offensive to some readers. Dog whistles are coded expressions with dual meanings: one intended for the general public (…
LLMs are Biased Teachers: Evaluating LLM Bias in Personalized Education
Iain Weissburg, Sathvika Anand, Sharon Levy +1
With the increasing adoption of large language models (LLMs) in education, concerns about inherent biases in these models have gained prominence. We evaluate LLMs for bias in the p…
Gender Bias in Decision-Making with Large Language Models: A Study of Relationship Conflicts
Sharon Levy, William D. Adler, Tahilin Sanchez Karver +2
Large language models (LLMs) acquire beliefs about gender from training data and can therefore generate text with stereotypical gender attitudes. Prior studies have demonstrated mo…