4 papers · 1 filter
Rate optimal learning of equilibria from data
Till Freihaut, Luca Viano, Emanuele Nevali +3
We close open theoretical gaps in Multi-Agent Imitation Learning (MAIL) by characterizing the limits of non-interactive MAIL and presenting the first interactive algorithm with nea…
ShiQ: Bringing back Bellman to LLMs
Pierre Clavier, Nathan Grinsztajn, Raphael Avalos +8
The fine-tuning of pre-trained large language models (LLMs) using reinforcement learning (RL) is generally formulated as direct policy optimization. This approach was naturally fav…
Learning Equilibria from Data: Provably Efficient Multi-Agent Imitation Learning
Till Freihaut, Luca Viano, Volkan Cevher +2
This paper provides the first expert sample complexity characterization for learning a Nash equilibrium from expert data in Markov Games. We show that a new quantity named the sing…
Solving robust MDPs as a sequence of static RL problems
Adil Zouitine, Matthieu Geist, Emmanuel Rachelson
Designing control policies whose performance level is guaranteed to remain above a given threshold in a span of environments is a critical feature for the adoption of reinforcement…