cs.CL2026
(Towards) Scalable Reliable Automated Evaluation with Large Language Models
Bertil Braun, Martin Forell
The paper presents a scalable framework for automatically evaluating large language model outputs using pairwise comparisons and an Elo rating system, achieving rankings that align…
#automated evaluation#large language models#pairwise comparison#elo rating