collaborators

9 papers

cs.AI2026

Recurrent Structural Policy Gradient for Partially Observable Mean Field Games

Clarisse Wibault, Johannes Forkel, Sebastian Towers +9

Mean Field Games (MFGs) provide a principled framework for modelling interactions in large population systems. However, algorithmic progress has been limited since model-free metho…

cs.AI2026

Asking the Right Questions: Improving Reasoning with Generated Stepping Stones

Hengyuan Hu, Tingchen Fu, Minqi Jiang +3

Recent years have witnessed tremendous progress in enabling LLMs to solve complex reasoning tasks such as math and coding. As we start to apply LLMs to harder tasks that they may n…

cs.LG2026

Learning to Reason at the Frontier of Learnability

Thomas Foster, Anya Sims, Johannes Forkel +2

Reinforcement learning is now widely adopted as the final stage of large language model training, especially for reasoning-style tasks such as maths problems. Typically, models att…

cs.LG2026

Infusion: Shaping Model Behavior by Editing Training Data via Influence Functions

J Rosser, Robert Kirk, Edward Grefenstette +2

Influence functions are commonly used to attribute model behavior to training documents. We explore the reverse: crafting training data that induces model behavior. Our framework,…

cs.LG2026

Improving Regret Approximation for Unsupervised Dynamic Environment Generation

Harry Mead, Bruno Lacerda, Jakob Foerster +1

Unsupervised Environment Design (UED) seeks to automatically generate training curricula for reinforcement learning (RL) agents, with the goal of improving generalisation and zero-…

cs.LG2026

Déjà Q: Open-Ended Evolution of Diverse, Learnable and Verifiable Problems

Willem Röpke, Samuel Coward, Andrei Lupu +3

Recent advances in reasoning models have yielded impressive results in mathematics and coding. However, most approaches rely on static datasets, which have been suggested to encour…