1 paper
Hiroyasu Usami, Keisuke Hara, Ayato Tsuboi +1
LLM-as-a-judge systems are now routinely used for open-ended model evaluation, where human preference annotation is costly, slow, and difficult to reproduce. Yet these judges are o…