164 citations · 368 across the 13 of their papers we have counts for
1 paper · 2 filters
Romain Froger, Pierre Andrews, Matteo Bettini +21
We introduce Gaia2, a benchmark for evaluating large language model agents in realistic, asynchronous environments. Unlike prior static or synchronous evaluations, Gaia2 introduces…