4 papers
The PokeAgent Challenge: Competitive and Long-Context Learning at Scale
Seth Karten, Jake Grigsby, Tersoo Upaa +28
We present the PokeAgent Challenge, a large-scale benchmark for decision-making research built on Pokemon's multi-agent battle system and expansive role-playing game (RPG) environm…
MUSIC: MUlti-Step Instruction Contrast for Multi-Turn Reward Models
Wenzhe Li, Shujian Zhang, Wenxuan Zhou +5
Evaluating the quality of multi-turn conversations is crucial for developing capable Large Language Models (LLMs), yet remains a significant challenge, often requiring costly human…
LLM Economist: Large Population Models and Mechanism Design in Multi-Agent Generative Simulacra
Seth Karten, Wenzhe Li, Zihan Ding +3
We present the LLM Economist, a novel framework that uses agent-based modeling to design and assess economic policies in strategic environments with hierarchical decision-making. A…
PokéChamp: an Expert-level Minimax Language Agent
Seth Karten, Andy Luu Nguyen, Chi Jin
We introduce PokéChamp, a minimax agent powered by Large Language Models (LLMs) for Pokémon battles. Built on a general framework for two-player competitive games, PokéChamp levera…