3 papers
cs.LG2026
Reward Models Inherit Value Biases from Pretraining
Brian Christian, Jessica A. F. Thompson, Elle Michelle Yang +4
Reward models (RMs) are central to aligning large language models (LLMs) with human values but have received less attention than pretrained and post-trained LLMs themselves. Becaus…
cs.CL2025
Reward Model Interpretability via Optimal and Pessimal Tokens
Brian Christian, Hannah Rose Kirk, Jessica A. F. Thompson +2
Reward modeling has emerged as a crucial component in aligning large language models with human values. Significant attention has focused on using reward models as a means for fine…
q-bio.NC2024
Zero-shot counting with a dual-stream neural network model
Jessica A. F. Thompson, Hannah Sheahan, Tsvetomira Dumbalska +3
Deep neural networks have provided a computational framework for understanding object recognition, grounded in the neurophysiology of the primate ventral stream, but fail to accoun…