Showing 2023Show all
2 papers · 1 filter
cs.LG2023
Towards Optimal Regret in Adversarial Linear MDPs with Bandit Feedback
Haolin Liu, Chen-Yu Wei, Julian Zimmert
We study online reinforcement learning in linear Markov decision processes with adversarial losses and bandit feedback, without prior knowledge on transitions or access to simulato…
cs.LG2023
Bypassing the Simulator: Near-Optimal Adversarial Linear Contextual Bandits
Haolin Liu, Chen-Yu Wei, Julian Zimmert
We consider the adversarial linear contextual bandit problem, where the loss vectors are selected fully adversarially and the per-round action set (i.e. the context) is drawn from…