3 citations · 5 across the 26 of their papers we have counts for
Showing 2025 · stat.MLShow all
2 papers · 2 filters
stat.ML2025
Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games
Antonio Ocello, Daniil Tiapkin, Lorenzo Mancini +2
We introduce Mean-Field Trust Region Policy Optimization (MF-TRPO), a novel algorithm designed to compute approximate Nash equilibria for ergodic Mean-Field Games (MFG) in finite s…
stat.ML2025
Proximal Point Nash Learning from Human Feedback
Daniil Tiapkin, Daniele Calandriello, Denis Belomestny +5
Traditional Reinforcement Learning from Human Feedback (RLHF) often relies on reward models, frequently assuming preference structures like the Bradley--Terry model, which may not…