3 papers
cs.LG2026
Robust AI Evaluation through Maximal Lotteries
Hadi Khalaf, Serena L. Wang, Daniel Halpern +3
The standard way to evaluate language models on subjective tasks is through pairwise comparisons: an annotator chooses the "better" of two responses to a prompt. Leaderboards aggre…
cs.LG2025
Inference-Time Reward Hacking in Large Language Models
Hadi Khalaf, Claudio Mayrink Verdun, Alex Oesterling +2
A common paradigm to improve the performance of large language models is optimizing for a reward model. Reward models assign a numerical score to an LLM's output that indicates, fo…
cs.AI2025
AI Alignment at Your Discretion
Maarten Buyl, Hadi Khalaf, Claudio Mayrink Verdun +3
In AI alignment, extensive latitude must be granted to annotators, either human or algorithmic, to judge which model outputs are `better' or `safer.' We refer to this latitude as a…