Showing cs.LGShow all
3 papers · 1 filter
cs.LG2025
A policy gradient approach for Finite Horizon Constrained Markov Decision Processes
Soumyajit Guin, Shalabh Bhatnagar
The infinite horizon setting is widely adopted for problems of reinforcement learning (RL). These invariably result in stationary policies that are optimal. In many situations, fin…
cs.LG2024
n-Step Temporal Difference Learning with Optimal n
Lakshmi Mandal, Shalabh Bhatnagar
We consider the problem of finding the optimal value of n in the n-step temporal difference (TD) learning algorithm. Our objective function for the optimization problem is the aver…
cs.LG2024
Actor-Critic or Critic-Actor? A Tale of Two Time Scales
Shalabh Bhatnagar, Vivek S. Borkar, Soumyajit Guin
We revisit the standard formulation of tabular actor-critic algorithm as a two time-scale stochastic approximation with value function computed on a faster time-scale and policy co…