12 citations · 56 across the 17 of their papers we have counts for
3 papers · 1 filter
Imagined Autocurricula
Ahmet H. Güzel, Matthew Thomas Jackson, Jarek Luca Liesen +4
Training agents to act in embodied environments typically requires vast training data or access to accurate simulation, neither of which exists for many cases in the real world. In…
Learning When to Plan: Efficiently Allocating Test-Time Compute for LLM Agents
Davide Paglieri, Bartłomiej Cupiał, Jonathan Cook +6
Training large language models (LLMs) to reason via reinforcement learning (RL) significantly improves their problem-solving capabilities. In agentic settings, existing methods lik…
BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
Davide Paglieri, Bartłomiej Cupiał, Samuel Coward +10
Large Language Models (LLMs) and Vision Language Models (VLMs) possess extensive knowledge and exhibit promising reasoning abilities, however, they still struggle to perform well i…