2 papers
cs.LG2023
Multi-Armed Bandits with Generalized Temporally-Partitioned Rewards
Ronald C. van den Broek, Rik Litjens, Tobias Sagis +3
Decision-making problems of sequential nature, where decisions made in the past may have an impact on the future, are used to model many practically important applications. In some…
cs.LG2022
Generalizing distribution of partial rewards for multi-armed bandits with temporally-partitioned rewards
Ronald C. van den Broek, Rik Litjens, Tobias Sagis +3
We investigate the Multi-Armed Bandit problem with Temporally-Partitioned Rewards (TP-MAB) setting in this paper. In the TP-MAB setting, an agent will receive subsets of the reward…