23 citations · 25 across the 2 of their papers we have counts for
2 papers
cs.CL2024★ 2 cited
Werewolf Arena: A Case Study in LLM Evaluation via Social Deduction
Suma Bailis, Jane Friedhoff, Feiyang Chen
This paper introduces Werewolf Arena, a novel framework for evaluating large language models (LLMs) through the lens of the classic social deduction game, Werewolf. In Werewolf Are…
cs.CL2022★ 23 cited
Dungeons and Dragons as a Dialog Challenge for Artificial Intelligence
Chris Callison-Burch, Gaurav Singh Tomar, Lara J. Martin +3
AI researchers have posited Dungeons and Dragons (D&D) as a challenge problem to test systems on various language-related capabilities. In this paper, we frame D&D specifically as…