1 paper
Diege Sun, Guanyi Chen, Zhao Fan +2
Large Language Models (LLMs) are increasingly used as automated evaluators in natural language generation, yet it remains unclear whether they can accurately replicate human judgme…