4 papers
Reward Models Inherit Value Biases from Pretraining
Brian Christian, Jessica A. F. Thompson, Elle Michelle Yang +4
Reward models (RMs) are central to aligning large language models (LLMs) with human values but have received less attention than pretrained and post-trained LLMs themselves. Becaus…
Reward Model Interpretability via Optimal and Pessimal Tokens
Brian Christian, Hannah Rose Kirk, Jessica A. F. Thompson +2
Reward modeling has emerged as a crucial component in aligning large language models with human values. Significant attention has focused on using reward models as a means for fine…
Self-Blinding and Counterfactual Self-Simulation Mitigate Biases and Sycophancy in Large Language Models
Brian Christian, Matan Mazor
Fair decisions require ignoring irrelevant, potentially biasing, information. To achieve this, decision-makers need to approximate what decision they would have made had they not k…
Personhood credentials: Artificial intelligence and the value of privacy-preserving tools to distinguish who is real online
Steven Adler, Zoë Hitzig, Shrey Jain +29
Anonymity is an important principle online. However, malicious actors have long used misleading identities to conduct fraud, spread disinformation, and carry out other deceptive sc…