1 paper
Weiyue Li, Minda Zhao, Weixuan Dong +12
Large language models (LLMs) are increasingly used as automated evaluators, yet prior works demonstrate that these LLM judges often lack consistency in scoring when the prompt is a…