activity
20172021
most citedModel-free Reinforcement Learning in Infinite-horizon Average-reward Markov Decision Processes

19 citations · 28 across the 3 of their papers we have counts for

collaborators

10 papers

cs.LG2021★ 3 cited

Online Learning for Cooperative Multi-Player Multi-Armed Bandits

William Chang, Mehdi Jafarnia-Jahromi, Rahul Jain

We introduce a framework for decentralized online learning for multi-armed bandits (MAB) with multiple cooperative players. The reward obtained by the players in each round depends…

cs.LG2021

A Bayesian Learning Algorithm for Unknown Zero-sum Stochastic Games with an Arbitrary Opponent

Mehdi Jafarnia-Jahromi, Rahul Jain, Ashutosh Nayyar

In this paper, we propose Posterior Sampling Reinforcement Learning for Zero-sum Stochastic Games (PSRL-ZSG), the first online learning algorithm that achieves Bayesian regret boun…

cs.LG2021★ 6 cited

Online Learning for Stochastic Shortest Path Model via Posterior Sampling

Mehdi Jafarnia-Jahromi, Liyu Chen, Rahul Jain +1

We consider the problem of online reinforcement learning for the Stochastic Shortest Path (SSP) problem modeled as an unknown MDP with an absorbing state. We propose PSRL-SSP, a si…

cs.LG2021

Implicit Finite-Horizon Approximation and Efficient Optimal Algorithms for Stochastic Shortest Path

Liyu Chen, Mehdi Jafarnia-Jahromi, Rahul Jain +1

We introduce a generic template for developing regret minimization algorithms in the Stochastic Shortest Path (SSP) model, which achieves minimax optimal regret as long as certain…

cs.LG2021

Online Learning for Unknown Partially Observable MDPs

Mehdi Jafarnia-Jahromi, Rahul Jain, Ashutosh Nayyar

Solving Partially Observable Markov Decision Processes (POMDPs) is hard. Learning optimal controllers for POMDPs when the model is unknown is harder. Online learning of optimal con…

cs.LG2020

Learning Infinite-horizon Average-reward MDPs with Linear Function Approximation

Chen-Yu Wei, Mehdi Jafarnia-Jahromi, Haipeng Luo +1

We develop several new algorithms for learning Markov Decision Processes in an infinite-horizon average-reward setting with linear function approximation. Using the optimism princi…