5 papers
Large Language Models Develop Novel Social Biases Through Adaptive Exploration
Addison J. Wu, Ryan Liu, Xuechunzi Bai +1
As large language models (LLMs) are adopted into frameworks that grant them the capacity to make real decisions, it is increasingly important to ensure that they are unbiased. In t…
Levels of Analysis for Large Language Models
Alexander Y. Ku, Declan Campbell, Xuechunzi Bai +10
Modern artificial intelligence systems, such as large language models, are increasingly powerful but also increasingly hard to understand. Recognizing this problem as analogous to…
Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race
Lihao Sun, Chengzhi Mao, Valentin Hofmann +1
Although value-aligned language models (LMs) appear unbiased in explicit bias evaluations, they often exhibit stereotypes in implicit word association tasks, raising concerns about…
Measuring Machine Learning Harms from Stereotypes Requires Understanding Who Is Harmed by Which Errors in What Ways
Angelina Wang, Xuechunzi Bai, Solon Barocas +1
As machine learning applications proliferate, we need an understanding of their potential for harm. However, current fairness metrics are rarely grounded in human psychological exp…
Measuring Implicit Bias in Explicitly Unbiased Large Language Models
Xuechunzi Bai, Angelina Wang, Ilia Sucholutsky +1
Large language models (LLMs) can pass explicit social bias tests but still harbor implicit biases, similar to humans who endorse egalitarian beliefs yet exhibit subtle biases. Meas…