11 papers
Sublogarithmic Swap Regret in Multiplayer General-Sum Games via Hybrid Regularization
Taira Tsuchiya
Swap regret governs the rate at which uncoupled learning dynamics converge to correlated equilibria in multiplayer general-sum games. Under full-information feedback, the best prev…
Scale-Invariant Fast Convergence in Games
Taira Tsuchiya, Haipeng Luo, Shinji Ito
Scale-invariance in games has recently emerged as a widely valued desirable property. Yet, almost all fast convergence guarantees in learning in games require prior knowledge of th…
Adversarial Learning in Games with Bandit Feedback: Logarithmic Pure-Strategy Maximin Regret
Shinji Ito, Haipeng Luo, Arnab Maiti +2
Learning to play zero-sum games is a fundamental problem in game theory and machine learning. While significant progress has been made in minimizing external regret in the self-pla…
Adapting to Stochastic and Adversarial Losses in Episodic MDPs with Aggregate Bandit Feedback
Shinji Ito, Kevin Jamieson, Haipeng Luo +2
We study online learning in finite-horizon episodic Markov decision processes (MDPs) under the challenging aggregate bandit feedback model, where the learner observes only the cumu…
Tight Regret Upper and Lower Bounds for Optimistic Hedge in Two-Player Zero-Sum Games
Taira Tsuchiya
In two-player zero-sum games, the learning dynamic based on optimistic Hedge achieves one of the best-known regret upper bounds among strongly-uncoupled learning dynamics. With an…
Reinforcement Learning from Adversarial Preferences in Tabular MDPs
Taira Tsuchiya, Shinji Ito, Haipeng Luo
We introduce a new framework of episodic tabular Markov decision processes (MDPs) with adversarial preferences, which we refer to as preference-based MDPs (PbMDPs). Unlike standard…