2 papers
cs.AI2025
Democratizing Diplomacy: A Harness for Evaluating Any Large Language Model on Full-Press Diplomacy
Alexander Duffy, Samuel J Paech, Ishana Shastri +4
We present the first evaluation harness that enables any out-of-the-box, local, Large Language Models (LLMs) to play full-press Diplomacy without fine-tuning or specialized trainin…
cs.AI2025
Strategic Intelligence in Large Language Models: Evidence from evolutionary Game Theory
Kenneth Payne, Baptiste Alloui-Cros
Are Large Language Models (LLMs) a new form of strategic intelligence, able to reason about goals in competitive settings? We present compelling supporting evidence. The Iterated P…