1 paper
Justin Zhao, Flor Miriam Plaza-del-Arco, Benjamin Genchel +1
As Large Language Models (LLMs) continue to evolve, evaluating them remains a persistent challenge. Many recent evaluations use LLMs as judges to score outputs from other LLMs, oft…