3 papers
cs.GT2025
Deviation Ratings: A General, Clone-Invariant Rating Method
Luke Marris, Siqi Liu, Ian Gemp +2
Many real-world multi-agent or multi-task evaluation scenarios can be naturally modelled as normal-form games due to inherent strategic (adversarial, cooperative, and mixed motive)…
cs.GT2025
Re-evaluating Open-ended Evaluation of Large Language Models
Siqi Liu, Ian Gemp, Luke Marris +3
Evaluation has traditionally focused on ranking candidates for a specific skill. Modern generalist models, such as Large Language Models (LLMs), decidedly outpace this paradigm. Op…
cs.GT2024
Convex Markov Games: A New Frontier for Multi-Agent Reinforcement Learning
Ian Gemp, Andreas Haupt, Luke Marris +2
Behavioral diversity, expert imitation, fairness, safety goals and others give rise to preferences in sequential decision making domains that do not decompose additively across tim…