1 paper
Ming-Zhe Dai, Chengxi Zhang
Consider a discounted Markov decision process with continuous action space in which, at each state visit, the controller draws a random pool of N candidate actions and selects am…