1 paper
Chen Peng, Di Zhang, Urbashi Mitra
In this paper, the causal bandit problem is investigated, with the objective of maximizing the long-term reward by selecting an optimal sequence of interventions on nodes in an unk…