2 papers
cs.CL2026
Multi-Expert Conformal Risk Control for Pairwise LLM Judging in Open-Ended Dialogue
Ming Cheng, Yusheng Dai, Qiuhong Ke +2
In this paper, we explore multi-expert Conformal Risk Control (CRC) algorithms for pairwise LLM-as-a-Judge evaluation in open-ended dialogue. Our core insight is that multi-expert…
cs.CL2026
Evaluating Language Models in Realistic Conversational Contexts
Ilija Subasic, Andrew Rabinovich, Zhao Chen
As Large Language Models (LLMs) are increasingly deployed to serve open-ended, multi-turn interactions, evaluating conversational quality at human scale has become a central challe…