1 paper
Haipeng Luo, Qingfeng Sun, Can Xu +6
Assessing the effectiveness of large language models (LLMs) presents substantial challenges. The method of conducting human-annotated battles in an online Chatbot Arena is a highly…