3 papers
cs.CL2026
The Dialect Tax: Dialectal Biases Persist throughout the Language Modeling Pipeline
Elle
Systematic dialectal performance gaps in language models (LMs) are well documented, but the source of these disparities within the modern language modeling pipeline remains unclear…
cs.LG2026
Reward Models Inherit Value Biases from Pretraining
Brian Christian, Jessica A. F. Thompson, Elle Michelle Yang +4
Reward models (RMs) are central to aligning large language models (LLMs) with human values but have received less attention than pretrained and post-trained LLMs themselves. Becaus…
cs.CL2025
Reward Model Perspectives: Whose Opinions Do Reward Models Reward?
Elle
Reward models (RMs) are central to the alignment of language models (LMs). An RM often serves as a proxy for human preferences to guide downstream LM behavior. However, our underst…