1 paper
Jun Wang, Jiamu Zhou, Muning Wen +7
Evaluating the performance of LLMs in multi-turn human-agent interactions presents significant challenges, particularly due to the complexity and variability of user behavior. In t…