1 paper
Ruomeng Ding, Yifei Pang, He Sun +3
Evaluation and alignment pipelines for large language models increasingly rely on LLM-based judges, whose behavior is guided by natural-language rubrics and validated on benchmarks…