1 paper · 1 filter
Xin Sun, Di Wu, Sijing Qin +3
Large language models (LLMs) are increasingly used as automated evaluators (LLM-as-a-Judge). This work challenges its reliability by showing that trust judgments by LLMs are biased…