39 citations · 77 across the 10 of their papers we have counts for
4 papers · 1 filter
Last-Iterate Convergent Policy Gradient Primal-Dual Methods for Constrained MDPs
Dongsheng Ding, Chen-Yu Wei, Kaiqing Zhang +1
We study the problem of computing an optimal policy of an infinite-horizon discounted constrained Markov decision process (constrained MDP). Despite the popularity of Lagrangian-ba…
First- and Second-Order Bounds for Adversarial Linear Contextual Bandits
Julia Olkhovskaya, Jack Mayo, Tim van Erven +2
We consider the adversarial linear contextual bandit setting, which allows for the loss functions associated with each of arms to change over time without restriction. Assuming…
No-Regret Online Reinforcement Learning with Adversarial Losses and Transitions
Tiancheng Jin, Junyan Liu, Chloé Rouyer +3
Existing online learning algorithms for adversarial Markov Decision Processes achieve regret after rounds of interactions even if the loss functions are chosen…
Uncoupled and Convergent Learning in Two-Player Zero-Sum Markov Games with Bandit Feedback
Yang Cai, Haipeng Luo, Chen-Yu Wei +1
We revisit the problem of learning in two-player zero-sum Markov games, focusing on developing an algorithm that is uncoupled, convergent, and rational, with non-asymptotic converg…