Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Expected Reward Prediction, with Applications to Model Routing
Kenan Hasanaliyev, Silas Alberti, Jenny Hamer +5
Reward models are a standard tool to score responses from LLMs. Reward models are built to rank responses to a fixed prompt sampled from a single model, for example to choose the b…
cs.CL2025
Bias in Language Models: Beyond Trick Tests and Toward RUTEd Evaluation
Kristian Lum, Jacy Reese Anthis, Kevin Robinson +2
Standard benchmarks of bias and fairness in large language models (LLMs) measure the association between the user attributes stated or implied by a prompt and the LLM's short text…
cs.CL2024
Transforming and Combining Rewards for Aligning Large Language Models
Zihao Wang, Chirag Nagpal, Jonathan Berant +4
A common approach for aligning language models to human preferences is to first learn a reward model from preference data, and then use this reward model to update the language mod…