4 papers
When is Routing Meaningful? Diversity and Robustness in Language Model Societies
Fantine Huot, Michael Kaisers, Mirella Lapata
Routing policies for multi-model systems are evaluated almost exclusively on task accuracy and inference cost. We argue that two properties, orthogonal to performance, determine wh…
Active Evaluation of General Agents: Problem Definition and Comparison of Baseline Algorithms
Marc Lanctot, Kate Larson, Ian Gemp +1
As intelligent agents become more generally-capable, i.e. able to master a wide variety of tasks, the complexity and cost of properly evaluating them rises significantly. Tasks tha…
Soft Condorcet Optimization for Ranking of General Agents
Marc Lanctot, Kate Larson, Michael Kaisers +7
Driving progress of AI models and agents requires comparing their performance on standardized benchmarks; for general agents, individual performances must be aggregated across a po…
Mastering Board Games by External and Internal Planning with Language Models
John Schultz, Jakub Adamek, Matej Jusup +13
Advancing planning and reasoning capabilities of Large Language Models (LLMs) is one of the key prerequisites towards unlocking their potential for performing reliably in complex a…