1 paper
Andreas Stephan, Dawei Zhu, Matthias AÃenmacher +2
To reduce the need for human annotations, large language models (LLMs) have been proposed as judges of the quality of other candidate models. The performance of LLM judges is typic…