1 paper
Liang Chen, Qi Liu, Wenhuan Lin +1
Multi-dimensional rubric-based dialogue evaluation is widely used to assess conversational AI, yet its criterion validity -- whether quality scores are associated with the downstre…