paper

When Can You Debias an LLM Judge? Identifiability Limits, a Test, and Designs for Top-k Ranking

arXiv:2607.02104

Abstract

Large language models (LLMs) are increasingly used as cheap, scalable judges that compare candidate outputs pairwise. Because such judges prefer verbose or well-formatted answers, the natural fix is to add bias covariates to a Bradley--Terry model and estimate the bias away. We show this cannot work as advertised: the quality/bias split is \emph{not identified} by pairwise comparisons, and the failure is exact -- across real judge-pools the profile likelihood over the coefficient is flat to \textbf{nats}, and scaling the comparisons buys none. A ``debiased'' score is selected by the prior, not recovered from data. Our contribution is accordingly not a better estimator but a characterization of \emph{when prior-based correction is justified}, plus designs that supply the missing information when it is not. The assumption the prior encodes -- quality is a priori uncorrelated with the covariate -- pays only while stays below a crossing point (configuration-dependent, --), which is what makes the same model help on LLMBar and hurt on SummEval and Nectar. We give two escapes: a \textbf{trusted-anchor gate} that decides per (judge, covariate, task) (no false enables in decisions at anchors, a rate our sample bounds at ), and a \textbf{paired rendering design}. Across fifteen real LLM judges bias is heterogeneous and capability-dependent: correction improves \topk{} recall by -- on five biased-but-competent cheap judges and is a no-op on frontier ones (Spearman between competence and gain over the competent judges, ), concentrating the benefit where at-scale evaluation happens.