2 citations · 4 across the 4 of their papers we have counts for
7 papers
Optimal cross-learning for contextual bandits with unknown context distributions
Jon Schneider, Julian Zimmert
We consider the problem of designing contextual bandit algorithms in the ``cross-learning'' setting of Balseiro et al., where the learner observes the loss for the action they play…
Towards Optimal Regret in Adversarial Linear MDPs with Bandit Feedback
Haolin Liu, Chen-Yu Wei, Julian Zimmert
We study online reinforcement learning in linear Markov decision processes with adversarial losses and bandit feedback, without prior knowledge on transitions or access to simulato…
Bypassing the Simulator: Near-Optimal Adversarial Linear Contextual Bandits
Haolin Liu, Chen-Yu Wei, Julian Zimmert
We consider the adversarial linear contextual bandit problem, where the loss vectors are selected fully adversarially and the per-round action set (i.e. the context) is drawn from…
A Blackbox Approach to Best of Both Worlds in Bandits and Beyond
Christoph Dann, Chen-Yu Wei, Julian Zimmert
Best-of-both-worlds algorithms for online learning which achieve near-optimal regret in both the adversarial and the stochastic regimes have received growing attention recently. Ex…
Best of Both Worlds Policy Optimization
Christoph Dann, Chen-Yu Wei, Julian Zimmert
Policy optimization methods are popular reinforcement learning algorithms in practice. Recent works have built theoretical foundation for them by proving regret bounds e…
A Provably Efficient Model-Free Posterior Sampling Method for Episodic Reinforcement Learning
Christoph Dann, Mehryar Mohri, Tong Zhang +1
Thompson Sampling is one of the most effective methods for contextual bandits and has been generalized to posterior sampling for certain MDP settings. However, existing posterior s…