3 papers
cs.LG2026
Direct Preference Optimization with Rating Information: Practical Algorithms and Provable Gains
Luca Viano, Ruida Zhou, Yifan Sun +4
The class of direct preference optimization (DPO) algorithms has emerged as a promising approach for solving the alignment problem in foundation models. These algorithms work with…
cs.LG2025
Rate optimal learning of equilibria from data
Till Freihaut, Luca Viano, Emanuele Nevali +3
We close open theoretical gaps in Multi-Agent Imitation Learning (MAIL) by characterizing the limits of non-interactive MAIL and presenting the first interactive algorithm with nea…
cs.GT2024
Polynomial Convergence of Bandit No-Regret Dynamics in Congestion Games
Leello Dadi, Ioannis Panageas, Stratis Skoulakis +2
We introduce an online learning algorithm in the bandit feedback model that, once adopted by all agents of a congestion game, results in game-dynamics that converge to an -appro…