2 papers
cs.LG2026
When LLM Judge Scores Look Good but Best-of-N Decisions Fail
Eddie Landesberg
Large language models are often used as judges to score candidate responses, then validated with a single global metric such as correlation with reference labels. This can be misle…
stat.ME2026
Causal Judge Evaluation: Calibrated Surrogate Metrics for LLM Systems
Eddie Landesberg, Manjari Narayan
Measuring long-run LLM outcomes (user satisfaction, expert judgment, downstream KPIs) is expensive. Teams default to cheap LLM judges, but uncalibrated proxies can invert rankings…