14 citations · 18 across the 9 of their papers we have counts for
9 papers
Contextual Multinomial Logit Bandits with General Value Functions
Mengxiao Zhang, Haipeng Luo
Contextual multinomial logit (MNL) bandits capture many real-world assortment recommendation problems such as online retailing/advertising. However, prior work has only considered…
Efficient Contextual Bandits with Uninformed Feedback Graphs
Mengxiao Zhang, Yuheng Zhang, Haipeng Luo +1
Bandits with feedback graphs are powerful online learning models that interpolate between the full information and classic bandit problems, capturing many real-life applications. A…
Near-Optimal Policy Optimization for Correlated Equilibrium in General-Sum Markov Games
Yang Cai, Haipeng Luo, Chen-Yu Wei +1
We study policy optimization algorithms for computing correlated equilibria in multi-player general-sum Markov Games. Previous results achieve convergence rate to a c…
Online Learning in Contextual Second-Price Pay-Per-Click Auctions
Mengxiao Zhang, Haipeng Luo
We study online learning in contextual pay-per-click auctions where at each of the rounds, the learner receives some context along with a set of ads and needs to make an estima…
Regret Matching+: (In)Stability and Fast Convergence in Games
Gabriele Farina, Julien Grand-Clément, Christian Kroer +2
Regret Matching+ (RM+) and its variants are important algorithms for solving large-scale games. However, a theoretical understanding of their success in practice is still a mystery…
Clairvoyant Regret Minimization: Equivalence with Nemirovski's Conceptual Prox Method and Extension to General Convex Games
Gabriele Farina, Christian Kroer, Chung-Wei Lee +1
A recent paper by Piliouras et al. [2021, 2022] introduces an uncoupled learning algorithm for normal-form games -- called Clairvoyant MWU (CMWU). In this note we show that CMWU is…