From the 1 of 9 linked papers with an AI index.
9 papers
Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning
Yang Xu, Swetha Ganesh, Vaneet Aggarwal
The paper develops and analyzes model‑free Q‑learning and actor‑critic algorithms that provably learn robust policies for infinite‑horizon average‑reward MDPs under various uncerta…
Bias-Controlled Primal-Dual Natural Actor-Critic: Optimal Rates for Constrained Multi-Objective Average-Reward RL
Ankur Naskar, Swetha Ganesh, Vaneet Aggarwal
Many reinforcement learning (RL) problems in the infinite-horizon average-reward setting require optimizing multiple conflicting objectives while satisfying multiple safety constra…
Adversary-Robust Learning from Fully Asynchronous Directional Derivative Estimates
Anik Kumar Paul, Nibedita Roy, Nagesh Talagani +3
We propose FAR-SIGN (Fully Asynchronous Robust optimization via SIGNed directional projections) for adversary-resilient learning in parameter-server--worker systems. FAR-SIGN achie…
Breaking the Bias Barrier in Concave Multi-Objective Reinforcement Learning
Swetha Ganesh, Vaneet Aggarwal
While standard reinforcement learning optimizes a single reward signal, many applications require optimizing a nonlinear utility over multiple objectives,…
Global Convergence for Average Reward Constrained MDPs with Primal-Dual Actor Critic Algorithm
Yang Xu, Swetha Ganesh, Washim Uddin Mondal +2
This paper investigates infinite-horizon average reward Constrained Markov Decision Processes (CMDPs) with general parametrization. We propose a Primal-Dual Natural Actor-Critic al…
Primal-Only Actor Critic Algorithm for Robust Constrained Average Cost MDPs
Anirudh Satheesh, Sooraj Sathish, Swetha Ganesh +2
In this work, we study the problem of finding robust and safe policies in Robust Constrained Average-Cost Markov Decision Processes (RCMDPs). A key challenge in this setting is the…