1 paper
Emmanuel Esposito, Saeed Masoudian, Hao Qiu +3
We study a K-armed bandit with delayed feedback and intermediate observations. We consider a model where intermediate observations have a form of a finite state, which is observe…