1 paper · 1 filter
Muhammad Qasim Elahi, Mahsa Ghasemi, Murat Kocaoglu
Causal knowledge about the relationships among decision variables and a reward variable in a bandit setting can accelerate the learning of an optimal decision. Current works often…