5 papers
What Drives Interactive Improvement from Feedback?
BartÅomiej CupiaÅ, Jan Åojek, MikoÅaj Garstecki +3
We study when natural-language feedback produces improvement beyond the gains obtainable from repeated attempts alone. In multi-turn language agent setting, higher final accuracy c…
Learning Multi-Agent Coordination via Sheaf-ADMM
Jeffrey Seely, BartÅomiej CupiaÅ, Llion Jones
We present a differentiable optimization framework for multi-agent coordination. An input is decomposed into overlapping local views, each processed by an agent that solves a conve…
Learning When to Plan: Efficiently Allocating Test-Time Compute for LLM Agents
Davide Paglieri, BartÅomiej CupiaÅ, Jonathan Cook +6
Training large language models (LLMs) to reason via reinforcement learning (RL) significantly improves their problem-solving capabilities. In agentic settings, existing methods lik…
BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
Davide Paglieri, BartÅomiej CupiaÅ, Samuel Coward +10
Large Language Models (LLMs) and Vision Language Models (VLMs) possess extensive knowledge and exhibit promising reasoning abilities, however, they still struggle to perform well i…
Fine-tuning Reinforcement Learning Models is Secretly a Forgetting Mitigation Problem
Maciej WoÅczyk, BartÅomiej CupiaÅ, Mateusz Ostaszewski +5
Fine-tuning is a widespread technique that allows practitioners to transfer pre-trained capabilities, as recently showcased by the successful applications of foundation models. How…