3 papers
cs.CL2026
Expected Reward Prediction, with Applications to Model Routing
Kenan Hasanaliyev, Silas Alberti, Jenny Hamer +5
Reward models are a standard tool to score responses from LLMs. Reward models are built to rank responses to a fixed prompt sampled from a single model, for example to choose the b…
cs.CL2025
Bias in Language Models: Beyond Trick Tests and Toward RUTEd Evaluation
Kristian Lum, Jacy Reese Anthis, Kevin Robinson +2
Standard benchmarks of bias and fairness in large language models (LLMs) measure the association between the user attributes stated or implied by a prompt and the LLM's short text…
cs.LG2025
Theoretical guarantees on the best-of-n alignment policy
Ahmad Beirami, Alekh Agarwal, Jonathan Berant +4
A simple and effective method for the inference-time alignment and scaling test-time compute of generative models is best-of- sampling, where samples are drawn from a refere…