3 papers
cs.LG2026
Reward Models Inherit Value Biases from Pretraining
Brian Christian, Jessica A. F. Thompson, Elle Michelle Yang +4
Reward models (RMs) are central to aligning large language models (LLMs) with human values but have received less attention than pretrained and post-trained LLMs themselves. Becaus…
cs.CL2026
Self-Blinding and Counterfactual Self-Simulation Mitigate Biases and Sycophancy in Large Language Models
Brian Christian, Matan Mazor
Fair decisions require ignoring irrelevant, potentially biasing, information. To achieve this, decision-makers need to approximate what decision they would have made had they not k…
cs.CL2025
Reward Model Interpretability via Optimal and Pessimal Tokens
Brian Christian, Hannah Rose Kirk, Jessica A. F. Thompson +2
Reward modeling has emerged as a crucial component in aligning large language models with human values. Significant attention has focused on using reward models as a means for fine…