collaborators

5 papers

cs.CL2026

Measuring and Mitigating Persona Distortions from AI Writing Assistance

Paul Röttger, Kobi Hackenburg, Hannah Rose Kirk +1

Hundreds of millions of people use artificial intelligence (AI) for writing assistance. Here, we evaluated how AI writing assistance distorts writer personas - their perceived beli…

cs.LG2026

Reward Models Inherit Value Biases from Pretraining

Brian Christian, Jessica A. F. Thompson, Elle Michelle Yang +4

Reward models (RMs) are central to aligning large language models (LLMs) with human values but have received less attention than pretrained and post-trained LLMs themselves. Becaus…

cs.CL2026

Reward Model Interpretability via Optimal and Pessimal Tokens

Brian Christian, Hannah Rose Kirk, Jessica A. F. Thompson +2

Reward modeling has emerged as a crucial component in aligning large language models with human values. Significant attention has focused on using reward models as a means for fine…

cs.CL2025

Measuring what Matters: Construct Validity in Large Language Model Benchmarks

Andrew M. Bean, Ryan Othniel Kearns, Angelika Romanou +39

Evaluating large language models (LLMs) is crucial for both assessing their capabilities and identifying safety or robustness issues prior to deployment. Reliably measuring abstrac…

cs.HC2025

Why human-AI relationships need socioaffective alignment

Hannah Rose Kirk, Iason Gabriel, Chris Summerfield +2

Humans strive to design safe AI systems that align with our goals and remain under our control. However, as AI capabilities advance, we face a new challenge: the emergence of deepe…