5 papers
The Horizon Threshold in Cooperative Multi-Agent Reward-Free Exploration
Idan Barnea, Orin Levy, Yishay Mansour
We study cooperative multi-agent reinforcement learning in the setting of reward-free exploration, where multiple agents jointly explore an unknown MDP in order to learn its dynami…
Near-Optimal Regret for Policy Optimization in Contextual MDPs with General Offline Function Approximation
Orin Levy, Aviv Rosenberg, Alon Cohen +1
We introduce \texttt{OPO-CMDP}, the first policy optimization algorithm for stochastic Contextual Markov Decision Process (CMDPs) under general offline function approximation. Our…
Optimal Regret for Policy Optimization in Contextual Bandits
Orin Levy, Yishay Mansour
We present the first high-probability optimal regret bound for a policy optimization technique applied to the problem of stochastic contextual multi-armed bandit (CMAB) with genera…
Regret Bounds for Adversarial Contextual Bandits with General Function Approximation and Delayed Feedback
Orin Levy, Liad Erez, Alon Cohen +1
We present regret minimization algorithms for the contextual multi-armed bandit (CMAB) problem over actions in the presence of delayed feedback, a scenario where loss observati…
Online Weighted Paging with Unknown Weights
Orin Levy, Noam Touitou, Aviv Rosenberg
Online paging is a fundamental problem in the field of online algorithms, in which one maintains a cache of slots as requests for fetching pages arrive online. In the weighted…