4 papers
Online Weighted Paging with Unknown Weights
Orin Levy, Noam Touitou, Aviv Rosenberg
Online paging is a fundamental problem in the field of online algorithms, in which one maintains a cache of slots as requests for fetching pages arrive online. In the weighted…
Warm-up Free Policy Optimization: Improved Regret in Linear Markov Decision Processes
Asaf Cassel, Aviv Rosenberg
Policy Optimization (PO) methods are among the most popular Reinforcement Learning (RL) algorithms in practice. Recently, Sherman et al. [2023a] proposed a PO-based algorithm with…
A Unified Analysis of Nonstochastic Delayed Feedback for Combinatorial Semi-Bandits, Linear Bandits, and MDPs
Dirk van der Hoeven, Lukas Zierahn, Tal Lancewicki +2
We derive a new analysis of Follow The Regularized Leader (FTRL) for online learning with delayed bandit feedback. By separating the cost of delayed feedback from that of bandit fe…
Delay-Adapted Policy Optimization and Improved Regret for Adversarial MDP with Delayed Bandit Feedback
Tal Lancewicki, Aviv Rosenberg, Dmitry Sotnikov
Policy Optimization (PO) is one of the most popular methods in Reinforcement Learning (RL). Thus, theoretical guarantees for PO algorithms have become especially important to the R…