1 paper
Gemma Zhang, Prachi Badarayani, Asmi Kumar +2
LLMs are increasingly used as automated judges for model training and evaluation, yet individual judges exhibit systematic biases that undermine reliability. Much of prior work has…