1 paper
Yuqi Tang, Kehua Feng, Yunfeng Wang +6
Evaluating the conversational abilities of large language models (LLMs) remains a challenging task. Current mainstream approaches primarily rely on the "LLM-as-a-judge" paradigm, w…