4 papers
Variance-Reduced Q-Learning over Static and Time-Varying Networks
Sreejeet Maity, Feng Zhu, Aritra Mitra +1
We investigate a decentralized reinforcement learning problem involving multiple agents that interact with the same Markov Decision Process (MDP). The agents can exchange informati…
Robust Asynchronous Q-Learning under Reward and State Corruption via Batching
Sreejeet Maity, Aritra Mitra
Motivated by reinforcement learning in harsh environments, we consider the problem of learning an optimal policy subject to adversarially corrupted feedback. Specifically, at each…
Corruption-Tolerant Asynchronous Q-Learning with Near-Optimal Rates
Sreejeet Maity, Aritra Mitra
We study the problem of learning the optimal policy in a discounted, infinite-horizon reinforcement learning (RL) setting in the presence of adversarially corrupted rewards. To add…
Adversarially-Robust TD Learning with Markovian Data: Finite-Time Rates and Fundamental Limits
Sreejeet Maity, Aritra Mitra
One of the most basic problems in reinforcement learning (RL) is policy evaluation: estimating the long-term return, i.e., value function, corresponding to a given fixed policy. Th…