2 papers
cs.LG2026
Operator-Theoretic Foundations and Policy Gradient Methods for General MDPs with Unbounded Costs
Abhishek Gupta, Aditya Mahajan
Markov decision processes (MDPs) is viewed as an optimization of an objective function over certain linear operators over general function spaces. A new existence result is establi…
math.OC2026
Sub-optimality bounds for certainty equivalent policies in partially observed systems
Berk Bozkurt, Aditya Mahajan, Ashutosh Nayyar +1
In this paper, we present a generalization of the certainty equivalence principle of stochastic control. One interpretation of the classical certainty equivalence principle for lin…