1 paper
Dongxu Lu, Johan Jeuring, Albert Gatt
Evaluating large language models (LLMs) in long-form, knowledge-grounded role-play dialogues remains challenging. This study compares LLM-generated and human-authored responses in…