2 papers
stat.ML2026
AURA: Adaptive Uncertainty-aware Refinement for LLM-as-a-Judge Auditing
Zilong Zhang, Yi-Ting Hung, Weiyi He +3
Large language models (LLMs) are increasingly used as judges for open-ended generation, as large-scale human evaluation is often expensive and difficult to scale, yet their prefere…
stat.ML2026
Quantifying and Auditing LLM Evaluation via Positive--Unlabeled Learning
Zilong Zhang, Yi-Ting Hung, Lei Ding +1
Large Language Models (LLMs) are increasingly used as judges for scalable evaluation, yet such LLM--as--a--Judge systems exhibit systematic biases that are decoupled from semantic…