1 paper · 1 filter
David Salinas, Omar Swelam, Frank Hutter
Evaluating Large Language Models (LLMs) often requires costly human annotations. To address this, LLM-based judges have been proposed, which compare the outputs of two LLMs enablin…