1 paper
Pengbin Feng, Chunlei Meng, Daozheng Qu +3
LLM judges are often evaluated with a single prompt and only a few repeated calls. When their verdicts vary, it remains unclear whether the variation comes from sampling noise with…