1 paper · 1 filter
Bertil Braun, Martin Forell
The paper presents a scalable framework for automatically evaluating large language model outputs using pairwise comparisons and an Elo rating system, achieving rankings that align…