3 papers
cs.LG2025
Near-optimal Regret Using Policy Optimization in Online MDPs with Aggregate Bandit Feedback
Tal Lancewicki, Yishay Mansour
We study online finite-horizon Markov Decision Processes with adversarially changing loss and aggregate bandit feedback (a.k.a full-bandit). Under this type of feedback, the agent…
cs.LG2025
Rising Rested MAB with Linear Drift
Omer Amichay, Yishay Mansour
We consider non-stationary multi-arm bandit (MAB) where the expected reward of each action follows a linear function of the number of times we executed the action. Our main result…
cs.LG2024
Batch Ensemble for Variance Dependent Regret in Stochastic Bandits
Asaf Cassel, Orin Levy, Yishay Mansour
Efficiently trading off exploration and exploitation is one of the key challenges in online Reinforcement Learning (RL). Most works achieve this by carefully estimating the model u…