1.1k citations · 1.6k across the 17 of their papers we have counts for
Showing 2025Show all
2 papers · 1 filter
cs.GT2025
Deviation Ratings: A General, Clone-Invariant Rating Method
Luke Marris, Siqi Liu, Ian Gemp +2
Many real-world multi-agent or multi-task evaluation scenarios can be naturally modelled as normal-form games due to inherent strategic (adversarial, cooperative, and mixed motive)…
cs.GT2025
Re-evaluating Open-ended Evaluation of Large Language Models
Siqi Liu, Ian Gemp, Luke Marris +3
Evaluation has traditionally focused on ranking candidates for a specific skill. Modern generalist models, such as Large Language Models (LLMs), decidedly outpace this paradigm. Op…