39 citations · 48 across the 4 of their papers we have counts for
10 papers · 1 filter
On Separation Between Best-Iterate, Random-Iterate, and Last-Iterate Convergence of Learning in Games
Yang Cai, Gabriele Farina, Julien Grand-Clément +4
Non-ergodic convergence of learning dynamics in games is widely studied recently because of its importance in both theory and practice. Recent work (Cai et al., 2024) showed that a…
Context-lumpable stochastic bandits
Chung-Wei Lee, Qinghua Liu, Yasin Abbasi-Yadkori +3
We consider a contextual bandit problem with contexts and actions. In each round , the learner observes a random context and chooses an action based on its pas…
Policy Optimization in Adversarial MDPs: Improved Exploration via Dilated Bonuses
Haipeng Luo, Chen-Yu Wei, Chung-Wei Lee
Policy optimization is a widely-used method in reinforcement learning. Due to its local-search nature, however, theoretical guarantees on global optimality often rely on extra assu…
Last-iterate Convergence in Extensive-Form Games
Chung-Wei Lee, Christian Kroer, Haipeng Luo
Regret-based algorithms are highly efficient at finding approximate Nash equilibria in sequential games such as poker games. However, most regret-based algorithms, including counte…
Achieving Near Instance-Optimality and Minimax-Optimality in Stochastic and Adversarial Linear Bandits Simultaneously
Chung-Wei Lee, Haipeng Luo, Chen-Yu Wei +2
In this work, we develop linear bandit algorithms that automatically adapt to different environments. By plugging a novel loss estimator into the optimization problem that characte…
Last-iterate Convergence of Decentralized Optimistic Gradient Descent/Ascent in Infinite-horizon Competitive Markov Games
Chen-Yu Wei, Chung-Wei Lee, Mengxiao Zhang +1
We study infinite-horizon discounted two-player zero-sum Markov games, and develop a decentralized algorithm that provably converges to the set of Nash equilibria under self-play.…