9 papers · 1 filter
Bias-Controlled Primal-Dual Natural Actor-Critic: Optimal Rates for Constrained Multi-Objective Average-Reward RL
Ankur Naskar, Swetha Ganesh, Vaneet Aggarwal
Many reinforcement learning (RL) problems in the infinite-horizon average-reward setting require optimizing multiple conflicting objectives while satisfying multiple safety constra…
Adversary-Robust Learning from Fully Asynchronous Directional Derivative Estimates
Anik Kumar Paul, Nibedita Roy, Nagesh Talagani +3
We propose FAR-SIGN (Fully Asynchronous Robust optimization via SIGNed directional projections) for adversary-resilient learning in parameter-server--worker systems. FAR-SIGN achie…
Breaking the Bias Barrier in Concave Multi-Objective Reinforcement Learning
Swetha Ganesh, Vaneet Aggarwal
While standard reinforcement learning optimizes a single reward signal, many applications require optimizing a nonlinear utility over multiple objectives, wh…
Primal-Only Actor Critic Algorithm for Robust Constrained Average Cost MDPs
Anirudh Satheesh, Sooraj Sathish, Swetha Ganesh +2
In this work, we study the problem of finding robust and safe policies in Robust Constrained Average-Cost Markov Decision Processes (RCMDPs). A key challenge in this setting is the…
Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning
Yang Xu, Swetha Ganesh, Vaneet Aggarwal
We study model-free methods for distributionally robust infinite-horizon average-reward Markov decision processes (MDPs). We present non-asymptotic convergence analyses of Q-learni…
Regret Analysis of Average-Reward Unichain MDPs via an Actor-Critic Approach
Swetha Ganesh, Vaneet Aggarwal
Actor-Critic methods are widely used for their scalability, yet existing theoretical guarantees for infinite-horizon average-reward Markov Decision Processes (MDPs) often rely on r…