1 paper
Ronald C. van den Broek, Rik Litjens, Tobias Sagis +3
We investigate the Multi-Armed Bandit problem with Temporally-Partitioned Rewards (TP-MAB) setting in this paper. In the TP-MAB setting, an agent will receive subsets of the reward…