2 papers
cs.LG2020
Lower Bounds for Policy Iteration on Multi-action MDPs
Kumar Ashutosh, Sarthak Consul, Bhishma Dedhia +3
Policy Iteration (PI) is a classical family of algorithms to compute an optimal policy for any given Markov Decision Problem (MDP). The basic idea in PI is to begin with some initi…
cs.LG2019
Analysis of Lower Bounds for Simple Policy Iteration
Sarthak Consul, Bhishma Dedhia, Kumar Ashutosh +1
Policy iteration is a family of algorithms that are used to find an optimal policy for a given Markov Decision Problem (MDP). Simple Policy iteration (SPI) is a type of policy iter…