Showing cs.LGShow all
3 papers · 1 filter
cs.LG2025
Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead
Uri Koren, Navdeep Kumar, Uri Gadot +3
Classical policy gradient (PG) methods in reinforcement learning frequently converge to suboptimal local optima, a challenge exacerbated in large or complex environments. This work…
cs.LG2024
On the Global Convergence of Policy Gradient in Average Reward Markov Decision Processes
Navdeep Kumar, Yashaswini Murthy, Itai Shufaro +3
We present the first finite time global convergence analysis of policy gradient in the context of infinite horizon average reward Markov decision processes (MDPs). Specifically, we…
cs.LG2023
An Efficient Solution to s-Rectangular Robust Markov Decision Processes
Navdeep Kumar, Kfir Levy, Kaixin Wang +1
We present an efficient robust value iteration for \texttt{s}-rectangular robust Markov Decision Processes (MDPs) with a time complexity comparable to standard (non-robust) MDPs wh…