1 paper · 1 filter
Hiroyasu Usami, Keisuke Hara, Ayato Tsuboi +1
LLM-as-a-judge systems are now routinely used for open-ended model evaluation, where human preference annotation is costly, slow, and difficult to reproduce. Yet these judges are o…