4 papers
To Nuke or Not to Nuke: LLMs' (Missing) Ethical Reasoning and Actions in a High-Stakes Decision-Making Simulation
John Chen, Sihan Cheng, Can Gurkan +1
Large language models (LLMs) are increasingly deployed as long-horizon agents with decision-making capacities. While LLMs can show ethical competence on dilemmas such as trolley pr…
Mutation Without Variation: Convergence Dynamics in LLM-Driven Program Evolution
Can Gurkan, Forrest Stonedahl, Uri Wilensky
When an LLM repeatedly mutates a program, does it explore new forms or circle back to the same ones? We study this question by analyzing LLM-driven mutation chains in the absence o…
CivBench: Progress-Based Evaluation for LLMs' Strategic Decision-Making in Civilization V
John Chen, Sihan Cheng, Can Gurkan +1
Evaluating strategic decision-making in LLM-based agents requires generative, competitive, and longitudinal environments, yet few benchmarks provide all three, and fewer still offe…
Vox Deorum: A Hybrid LLM Architecture for 4X / Grand Strategy Game AI -- Lessons from Civilization V
John Chen, Sihan Cheng, Can Gurkan +2
Large Language Models' capacity to reason in natural language makes them uniquely promising for 4X and grand strategy games, enabling more natural human-AI gameplay interactions su…