1 paper
Deeksha Sinha, Karthik Abinav Sankararama, Abbas Kazerouni +1
In this paper, we consider a novel variant of the multi-armed bandit (MAB) problem, MAB with cost subsidy, which models many real-life applications where the learning agent has to…