activity
20172020
most citedA New Algorithm for Non-stationary Contextual Bandits: Efficient, Optimal, and Parameter-free

39 citations · 75 across the 5 of their papers we have counts for

collaborators

7 papers

cs.LG2020

Federated Residual Learning

Alekh Agarwal, John Langford, Chen-Yu Wei

We study a new form of federated learning where the clients train personalized local models and make predictions jointly with the server-side shared model. Using this new federated…

cs.LG201919 cited

Model-free Reinforcement Learning in Infinite-horizon Average-reward Markov Decision Processes

Chen-Yu Wei, Mehdi Jafarnia-Jahromi, Haipeng Luo +2

Model-free reinforcement learning is known to be memory and computation efficient and more amendable to large scale problems. In this paper, two model-free algorithms are introduce…

cs.LG20193 cited

Bandit Multiclass Linear Classification: Efficient Algorithms for the Separable Case

Alina Beygelzimer, Dávid Pál, Balázs Szörényi +3

We study the problem of efficient online multiclass linear classification with bandit feedback, where all examples belong to one of classes and lie in the -dimensional Eucli…

cs.LG201939 cited

A New Algorithm for Non-stationary Contextual Bandits: Efficient, Optimal, and Parameter-free

Yifang Chen, Chung-Wei Lee, Haipeng Luo +1

We propose the first contextual bandit algorithm that is parameter-free, efficient, and optimal in terms of dynamic regret. Specifically, our algorithm achieves dynamic regret $\ma…

cs.LG20197 cited

Improved Path-length Regret Bounds for Bandits

Sébastien Bubeck, Yuanzhi Li, Haipeng Luo +1

We study adaptive regret bounds in terms of the variation of the losses (the so-called path-length bounds) for both multi-armed bandit and more generally linear bandit. We first sh…

cs.LG2019

Beating Stochastic and Adversarial Semi-bandits Optimally and Simultaneously

Julian Zimmert, Haipeng Luo, Chen-Yu Wei

We develop the first general semi-bandit algorithm that simultaneously achieves regret for stochastic environments and regret for adve…