1 paper · 1 filter
Suma Bailis, Jane Friedhoff, Feiyang Chen
This paper introduces Werewolf Arena, a novel framework for evaluating large language models (LLMs) through the lens of the classic social deduction game, Werewolf. In Werewolf Are…