1 paper
Daniil Tiapkin, Evgenii Chzhen, Gilles Stoltz
We consider the problem of learning in adversarial Markov decision processes [MDPs] with an oblivious adversary in a full-information setting. The agent interacts with an environme…