1 citations · 1 across the 3 of their papers we have counts for
Showing stat.MLShow all
2 papers · 1 filter
stat.ML2025
Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games
Antonio Ocello, Daniil Tiapkin, Lorenzo Mancini +2
We introduce Mean-Field Trust Region Policy Optimization (MF-TRPO), a novel algorithm designed to compute approximate Nash equilibria for ergodic Mean-Field Games (MFG) in finite s…
stat.ML2022
Optimistic Posterior Sampling for Reinforcement Learning with Few Samples and Tight Guarantees
Daniil Tiapkin, Denis Belomestny, Daniele Calandriello +6
We consider reinforcement learning in an environment modeled by an episodic, finite, stage-dependent Markov decision process of horizon with states, and actions. The pe…