1 paper
Zilong Zhang, Yi-Ting Hung, Weiyi He +3
Large language models (LLMs) are increasingly used as judges for open-ended generation, as large-scale human evaluation is often expensive and difficult to scale, yet their prefere…