3 citations · 3 across the 2 of their papers we have counts for
3 papers
Lower Bounds for Policy Iteration on Multi-action MDPs
Kumar Ashutosh, Sarthak Consul, Bhishma Dedhia +3
Policy Iteration (PI) is a classical family of algorithms to compute an optimal policy for any given Markov Decision Problem (MDP). The basic idea in PI is to begin with some initi…
Bandit algorithms: Letting go of logarithmic regret for statistical robustness
Kumar Ashutosh, Jayakrishnan Nair, Anmol Kagrecha +1
We study regret minimization in a stochastic multi-armed bandit setting and establish a fundamental trade-off between the regret suffered under an algorithm, and its statistical ro…
Analysis of Lower Bounds for Simple Policy Iteration
Sarthak Consul, Bhishma Dedhia, Kumar Ashutosh +1
Policy iteration is a family of algorithms that are used to find an optimal policy for a given Markov Decision Problem (MDP). Simple Policy iteration (SPI) is a type of policy iter…