1 paper
Ziyi Zhu, Olivier Tieleman, Alexey Bukhtiyarov +1
LLM-as-judge evaluation has become standard practice for open-ended model assessment; however, judges exhibit systematic biases that cannot be averaged out by increasing the number…