2 papers
cs.LG2025
Regret Bounds for Adversarial Contextual Bandits with General Function Approximation and Delayed Feedback
Orin Levy, Liad Erez, Alon Cohen +1
We present regret minimization algorithms for the contextual multi-armed bandit (CMAB) problem over actions in the presence of delayed feedback, a scenario where loss observati…
cs.LG2024
Batch Ensemble for Variance Dependent Regret in Stochastic Bandits
Asaf Cassel, Orin Levy, Yishay Mansour
Efficiently trading off exploration and exploitation is one of the key challenges in online Reinforcement Learning (RL). Most works achieve this by carefully estimating the model u…