4 citations · 7 across the 5 of their papers we have counts for
8 papers
Match or Replay: Self Imitating Proximal Policy Optimization
Gaurav Chaudhary, Laxmidhar Behera, Washim Uddin Mondal
Reinforcement Learning (RL) agents often struggle with inefficient exploration, particularly in environments with sparse rewards. Traditional exploration strategies can lead to slo…
Global Convergence of Average Reward Constrained MDPs with Neural Critic and General Policy Parameterization
Anirudh Satheesh, Pankaj Kumar Barman, Washim Uddin Mondal +1
We study infinite-horizon Constrained Markov Decision Processes (CMDPs) with general policy parameterizations and multi-layer neural network critics. Existing theoretical analyses…
MOORL: A Framework for Integrating Offline-Online Reinforcement Learning
Gaurav Chaudhary, Wassim Uddin Mondal, Laxmidhar Behera
Sample efficiency and exploration remain critical challenges in Deep Reinforcement Learning (DRL), particularly in complex domains. Offline RL, which enables agents to learn optima…
Global Convergence for Average Reward Constrained MDPs with Primal-Dual Actor Critic Algorithm
Yang Xu, Swetha Ganesh, Washim Uddin Mondal +2
This paper investigates infinite-horizon average reward Constrained Markov Decision Processes (CMDPs) with general parametrization. We propose a Primal-Dual Natural Actor-Critic al…
Finite-Sample Analysis of Policy Evaluation for Robust Average Reward Reinforcement Learning
Yang Xu, Washim Uddin Mondal, Vaneet Aggarwal
We present the first finite-sample analysis of policy evaluation in robust average-reward Markov Decision Processes (MDPs). Prior work in this setting have established only asympto…
On the Near-Optimality of Local Policies in Large Cooperative Multi-Agent Reinforcement Learning
Washim Uddin Mondal, Vaneet Aggarwal, Satish V. Ukkusuri
We show that in a cooperative -agent network, one can design locally executable policies for the agents such that the resulting discounted sum of average rewards (value) well ap…