1 paper
Xinyu Li, Yi Zhou, Guanqun Cao +3
Large language models (LLMs) are increasingly used as evaluators to assess output quality and preference alignment, yet providing reliable guarantees of agreement with human judgme…