1 paper · 1 filter
Romain Froger, Pierre Andrews, Matteo Bettini +21
We introduce Gaia2, a benchmark for evaluating large language model agents in realistic, asynchronous environments. Unlike prior static or synchronous evaluations, Gaia2 introduces…