1 paper
Jitian Zhao, Changho Shin, Tzu-Heng Huang +2
LLM-as-a-judge ensembles are the standard paradigm for scalable evaluation, but their aggregation mechanisms suffer from a fundamental flaw: they implicitly assume that judges prov…