1 paper
Ilija Subasic, Andrew Rabinovich, Zhao Chen
As Large Language Models (LLMs) are increasingly deployed to serve open-ended, multi-turn interactions, evaluating conversational quality at human scale has become a central challe…