14 citations · 14 across the 2 of their papers we have counts for
2 papers
math.OC2026
Sub-optimality bounds for certainty equivalent policies in partially observed systems
Berk Bozkurt, Aditya Mahajan, Ashutosh Nayyar +1
In this paper, we present a generalization of the certainty equivalence principle of stochastic control. One interpretation of the classical certainty equivalence principle for lin…
cs.LG2017★ 14 cited
Learning Unknown Markov Decision Processes: A Thompson Sampling Approach
Yi Ouyang, Mukul Gagrani, Ashutosh Nayyar +1
We consider the problem of learning an unknown Markov Decision Process (MDP) that is weakly communicating in the infinite horizon setting. We propose a Thompson Sampling-based rein…