1 paper · 1 filter
Brandon Trabucco, Albert Qu, Simon Li +1
This paper investigates methods for estimating the optimal stochastic control policy for a Markov Decision Process with unknown transition dynamics and an unknown reward function.…