3 papers
cs.AI2026
Gaia2: Benchmarking LLM Agents on Dynamic and Asynchronous Environments
Romain Froger, Pierre Andrews, Matteo Bettini +21
We introduce Gaia2, a benchmark for evaluating large language model agents in realistic, asynchronous environments. Unlike prior static or synchronous evaluations, Gaia2 introduces…
cs.AI2025
ARE: Scaling Up Agent Environments and Evaluations
Romain Froger, Pierre Andrews, Matteo Bettini +21
We introduce Meta Agents Research Environments (ARE), a research platform for scalable creation of environments, integration of synthetic or real applications, and execution of age…
cs.LG2024
Optimal Design for Reward Modeling in RLHF
Antoine Scheid, Etienne Boursier, Alain Durmus +4
Reinforcement Learning from Human Feedback (RLHF) has become a popular approach to align language models (LMs) with human preferences. This method involves collecting a large datas…