activity
20152021
most citedA New Algorithm for Non-stationary Contextual Bandits: Efficient, Optimal, and Parameter-free

39 citations · 124 across the 12 of their papers we have counts for

collaborators

43 papers

cs.LG2022

Near-Optimal Goal-Oriented Reinforcement Learning in Non-Stationary Environments

Liyu Chen, Haipeng Luo

We initiate the study of dynamic regret minimization for goal-oriented reinforcement learning modeled by a non-stationary stochastic shortest path problem with changing cost and tr…

cs.LG2022

Corralling a Larger Band of Bandits: A Case Study on Switching Regret for Linear Bandits

Haipeng Luo, Mengxiao Zhang, Peng Zhao +1

We consider the problem of combining and learning over a set of adversarial bandit algorithms with the goal of adaptively tracking the best one on the fly. The CORRAL algorithm of…

cs.LG20221 cited

Adaptive Bandit Convex Optimization with Heterogeneous Curvature

Haipeng Luo, Mengxiao Zhang, Peng Zhao

We consider the problem of adversarial bandit convex optimization, that is, online learning over a sequence of arbitrary convex loss functions with only one function evaluation for…

cs.LG2022

Policy Optimization for Stochastic Shortest Path

Liyu Chen, Haipeng Luo, Aviv Rosenberg

Policy optimization is among the most popular and successful reinforcement learning algorithms, and there is increasing interest in understanding its theoretical guarantees. In thi…

cs.GT20223 cited

Kernelized Multiplicative Weights for 0/1-Polyhedral Games: Bridging the Gap Between Learning in Extensive-Form and Normal-Form Games

Gabriele Farina, Chung-Wei Lee, Haipeng Luo +1

While extensive-form games (EFGs) can be converted into normal-form games (NFGs), doing so comes at the cost of an exponential blowup of the strategy space. So, progress on NFGs an…

cs.LG20223 cited

Learning Infinite-Horizon Average-Reward Markov Decision Processes with Constraints

Liyu Chen, Rahul Jain, Haipeng Luo

We study regret minimization for infinite-horizon average-reward Markov Decision Processes (MDPs) under cost constraints. We start by designing a policy optimization algorithm with…