39 citations · 75 across the 5 of their papers we have counts for
7 papers
Federated Residual Learning
Alekh Agarwal, John Langford, Chen-Yu Wei
We study a new form of federated learning where the clients train personalized local models and make predictions jointly with the server-side shared model. Using this new federated…
Model-free Reinforcement Learning in Infinite-horizon Average-reward Markov Decision Processes
Chen-Yu Wei, Mehdi Jafarnia-Jahromi, Haipeng Luo +2
Model-free reinforcement learning is known to be memory and computation efficient and more amendable to large scale problems. In this paper, two model-free algorithms are introduce…
Bandit Multiclass Linear Classification: Efficient Algorithms for the Separable Case
Alina Beygelzimer, Dávid Pál, Balázs Szörényi +3
We study the problem of efficient online multiclass linear classification with bandit feedback, where all examples belong to one of classes and lie in the -dimensional Eucli…
A New Algorithm for Non-stationary Contextual Bandits: Efficient, Optimal, and Parameter-free
Yifang Chen, Chung-Wei Lee, Haipeng Luo +1
We propose the first contextual bandit algorithm that is parameter-free, efficient, and optimal in terms of dynamic regret. Specifically, our algorithm achieves dynamic regret $\ma…
Improved Path-length Regret Bounds for Bandits
Sébastien Bubeck, Yuanzhi Li, Haipeng Luo +1
We study adaptive regret bounds in terms of the variation of the losses (the so-called path-length bounds) for both multi-armed bandit and more generally linear bandit. We first sh…
Beating Stochastic and Adversarial Semi-bandits Optimally and Simultaneously
Julian Zimmert, Haipeng Luo, Chen-Yu Wei
We develop the first general semi-bandit algorithm that simultaneously achieves regret for stochastic environments and regret for adve…