2 papers
cs.AI2025
Efficient Computation of Blackwell Optimal Policies using Rational Functions
Dibyangshu Mukherjee, Shivaram Kalyanakrishnan
Markov Decision Problems (MDPs) provide a foundational framework for modelling sequential decision-making across diverse domains, guided by optimality criteria such as discounted a…
cs.AI2025
Howard's Policy Iteration is Subexponential for Deterministic Markov Decision Problems with Rewards of Fixed Bit-size and Arbitrary Discount Factor
Dibyangshu Mukherjee, Shivaram Kalyanakrishnan
Howard's Policy Iteration (HPI) is a classic algorithm for solving Markov Decision Problems (MDPs). HPI uses a "greedy" switching rule to update from any non-optimal policy to a do…