1 paper
Maike Züfle, Patrícia Schmidtová, Vilém Zouhar +8
Measuring how successful a conversation is remains difficult, even for humans judging spoken dialogue. We evaluate state-of-the-art LLMs as pointwise and pairwise judges of convers…