1 paper
Jasmine Qi, Danylo Dantsev, Muyang Sun
LLM-as-Judge systems are widely deployed for automated evaluation, yet practitioners lack reliable methods to know when a judge's verdict should be trusted. Token log-probabilities…