most citedA Provably Efficient Model-Free Posterior Sampling Method for Episodic Reinforcement Learning

2 citations · 4 across the 4 of their papers we have counts for

collaborators

7 papers

cs.LG2024

Optimal cross-learning for contextual bandits with unknown context distributions

Jon Schneider, Julian Zimmert

We consider the problem of designing contextual bandit algorithms in the ``cross-learning'' setting of Balseiro et al., where the learner observes the loss for the action they play…

cs.LG2023

Towards Optimal Regret in Adversarial Linear MDPs with Bandit Feedback

Haolin Liu, Chen-Yu Wei, Julian Zimmert

We study online reinforcement learning in linear Markov decision processes with adversarial losses and bandit feedback, without prior knowledge on transitions or access to simulato…

cs.LG2023

Bypassing the Simulator: Near-Optimal Adversarial Linear Contextual Bandits

Haolin Liu, Chen-Yu Wei, Julian Zimmert

We consider the adversarial linear contextual bandit problem, where the loss vectors are selected fully adversarially and the per-round action set (i.e. the context) is drawn from…

cs.LG2023

A Blackbox Approach to Best of Both Worlds in Bandits and Beyond

Christoph Dann, Chen-Yu Wei, Julian Zimmert

Best-of-both-worlds algorithms for online learning which achieve near-optimal regret in both the adversarial and the stochastic regimes have received growing attention recently. Ex…

cs.LG2023

Best of Both Worlds Policy Optimization

Christoph Dann, Chen-Yu Wei, Julian Zimmert

Policy optimization methods are popular reinforcement learning algorithms in practice. Recent works have built theoretical foundation for them by proving regret bounds e…

cs.LG20222 cited

A Provably Efficient Model-Free Posterior Sampling Method for Episodic Reinforcement Learning

Christoph Dann, Mehryar Mohri, Tong Zhang +1

Thompson Sampling is one of the most effective methods for contextual bandits and has been generalized to posterior sampling for certain MDP settings. However, existing posterior s…