Online Convex Optimization in Adversarial Markov Decision Processes
arXiv:1905.07773
Abstract
We consider online learning in episodic loop-free Markov decision processes (MDPs), where the loss function can change arbitrarily between episodes, and the transition function is not known to the learner. We show regret bound, where is the number of episodes, is the state space, is the action space, and is the length of each episode. Our online algorithm is implemented using entropic regularization methodology, which allows to extend the original adversarial MDP model to handle convex performance criteria (different ways to aggregate the losses of a single episode) , as well as improve previous regret bounds.
Cited by in corpus (8)
- Exploration-Exploitation in Constrained MDPs
- Reinforcement Learning for Non-Stationary Markov Decision Processes: The Blessing of (More) Optimism
- Improved Corruption Robust Algorithms for Episodic Reinforcement Learning
- Policy Optimization in Adversarial MDPs: Improved Exploration via Dilated Bonuses
- Model-Free Learning for Two-Player Zero-Sum Partially Observable Markov Games with Perfect Recall
- The best of both worlds: stochastic and adversarial episodic MDPs with unknown transition
- Refined Analysis of FPL for Adversarial Markov Decision Processes
- Adaptive Sampling for Estimating Distributions: A Bayesian Upper Confidence Bound Approach