activity
20242026
collaborators

5 papers

cs.LG2026

The Horizon Threshold in Cooperative Multi-Agent Reward-Free Exploration

Idan Barnea, Orin Levy, Yishay Mansour

We study cooperative multi-agent reinforcement learning in the setting of reward-free exploration, where multiple agents jointly explore an unknown MDP in order to learn its dynami…

cs.LG2026

Near-Optimal Regret for Policy Optimization in Contextual MDPs with General Offline Function Approximation

Orin Levy, Aviv Rosenberg, Alon Cohen +1

We introduce \texttt{OPO-CMDP}, the first policy optimization algorithm for stochastic Contextual Markov Decision Process (CMDPs) under general offline function approximation. Our…

cs.LG2026

Optimal Regret for Policy Optimization in Contextual Bandits

Orin Levy, Yishay Mansour

We present the first high-probability optimal regret bound for a policy optimization technique applied to the problem of stochastic contextual multi-armed bandit (CMAB) with genera…

cs.LG2025

Regret Bounds for Adversarial Contextual Bandits with General Function Approximation and Delayed Feedback

Orin Levy, Liad Erez, Alon Cohen +1

We present regret minimization algorithms for the contextual multi-armed bandit (CMAB) problem over actions in the presence of delayed feedback, a scenario where loss observati…

cs.LG2024

Online Weighted Paging with Unknown Weights

Orin Levy, Noam Touitou, Aviv Rosenberg

Online paging is a fundamental problem in the field of online algorithms, in which one maintains a cache of slots as requests for fetching pages arrive online. In the weighted…