15 citations · 22 across the 24 of their papers we have counts for
1 paper · 1 filter
Felipe Maia Polo, Xinhe Wang, Mikhail Yurochkin +3
Large language models are increasingly used as judges (LLM-as-a-judge) to evaluate model outputs at scale, but their assessments often diverge systematically from human judgments.…