1 paper
Brandon Trabucco, Albert Qu, Simon Li +1
This paper investigates methods for estimating the optimal stochastic control policy for a Markov Decision Process with unknown transition dynamics and an unknown reward function.…