1 paper · 1 filter
Marc Lanctot, Kate Larson, Michael Kaisers +7
Driving progress of AI models and agents requires comparing their performance on standardized benchmarks; for general agents, individual performances must be aggregated across a po…