1 paper
Nicolas Gast, Dheeraj Narasimha
We consider the discrete time infinite horizon average reward restless markovian bandit (RMAB) problem. We propose a \emph{model predictive control} based non-stationary policy wit…