activity
20242026
collaborators

5 papers

cs.CY2026

Large Language Models Develop Novel Social Biases Through Adaptive Exploration

Addison J. Wu, Ryan Liu, Xuechunzi Bai +1

As large language models (LLMs) are adopted into frameworks that grant them the capacity to make real decisions, it is increasingly important to ensure that they are unbiased. In t…

cs.CL2026

Levels of Analysis for Large Language Models

Alexander Y. Ku, Declan Campbell, Xuechunzi Bai +10

Modern artificial intelligence systems, such as large language models, are increasingly powerful but also increasingly hard to understand. Recognizing this problem as analogous to…

cs.CL2025

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race

Lihao Sun, Chengzhi Mao, Valentin Hofmann +1

Although value-aligned language models (LMs) appear unbiased in explicit bias evaluations, they often exhibit stereotypes in implicit word association tasks, raising concerns about…

cs.CY2025

Measuring Machine Learning Harms from Stereotypes Requires Understanding Who Is Harmed by Which Errors in What Ways

Angelina Wang, Xuechunzi Bai, Solon Barocas +1

As machine learning applications proliferate, we need an understanding of their potential for harm. However, current fairness metrics are rarely grounded in human psychological exp…

cs.CY2024

Measuring Implicit Bias in Explicitly Unbiased Large Language Models

Xuechunzi Bai, Angelina Wang, Ilia Sucholutsky +1

Large language models (LLMs) can pass explicit social bias tests but still harbor implicit biases, similar to humans who endorse egalitarian beliefs yet exhibit subtle biases. Meas…