activity
20242026
collaborators

6 papers

cs.AI2026

Evaluating Implicit Biases in LLM Reasoning through Logic Grid Puzzles

Fatima Jahara, Mark Dredze, Sharon Levy

While recent safety guardrails effectively suppress overtly biased outputs, subtler forms of social bias emerge during complex logical reasoning tasks that evade current evaluation…

cs.CL2026

Knowing But Not Doing: Convergent Morality and Divergent Action in LLMs

Jen-tse Huang, Jiantong Qin, Xueli Qiu +3

Value alignment is central to the development of safe and socially compatible artificial intelligence. However, how Large Language Models (LLMs) represent and enact human values in…

cs.CL2025

Characterizing Selective Refusal Bias in Large Language Models

Adel Khorramrouz, Sharon Levy

Safety guardrails in large language models(LLMs) are developed to prevent malicious users from generating toxic content at a large scale. However, these measures can inadvertently…

cs.CL2025

Making FETCH! Happen: Finding Emergent Dog Whistles Through Common Habitats

Kuleen Sasse, Carlos Aguirre, Isabel Cachola +2

WARNING: This paper contains content that maybe upsetting or offensive to some readers. Dog whistles are coded expressions with dual meanings: one intended for the general public (…

cs.CL2025

LLMs are Biased Teachers: Evaluating LLM Bias in Personalized Education

Iain Weissburg, Sathvika Anand, Sharon Levy +1

With the increasing adoption of large language models (LLMs) in education, concerns about inherent biases in these models have gained prominence. We evaluate LLMs for bias in the p…

cs.CL2024

Gender Bias in Decision-Making with Large Language Models: A Study of Relationship Conflicts

Sharon Levy, William D. Adler, Tahilin Sanchez Karver +2

Large language models (LLMs) acquire beliefs about gender from training data and can therefore generate text with stereotypical gender attitudes. Prior studies have demonstrated mo…