activity
20182021
most citedConstrained Markov Decision Processes via Backward Value Functions

16 citations · 26 across the 3 of their papers we have counts for

collaborators

6 papers

cs.LG20215 cited

A Survey of Exploration Methods in Reinforcement Learning

Susan Amin, Maziar Gomrokchi, Harsh Satija +2

Exploration is an essential component of reinforcement learning algorithms, where agents need to learn how to predict and control unknown and often stochastic environments. Reinfor…

cs.LG20215 cited

Multi-Objective SPIBB: Seldonian Offline Policy Improvement with Safety Constraints in Finite MDPs

Harsh Satija, Philip S. Thomas, Joelle Pineau +1

We study the problem of Safe Policy Improvement (SPI) under constraints in the offline Reinforcement Learning (RL) setting. We consider the scenario where: (i) we have a dataset co…

cs.LG2020

Locally Persistent Exploration in Continuous Control Tasks with Sparse Rewards

Susan Amin, Maziar Gomrokchi, Hossein Aboutalebi +2

A major challenge in reinforcement learning is the design of exploration strategies, especially for environments with sparse reward structures and continuous state and action space…

cs.LG202016 cited

Constrained Markov Decision Processes via Backward Value Functions

Harsh Satija, Philip Amortila, Joelle Pineau

Although Reinforcement Learning (RL) algorithms have found tremendous success in simulated domains, they often cannot directly be applied to physical systems, especially in cases w…

cs.LG2018

Randomized Value Functions via Multiplicative Normalizing Flows

Ahmed Touati, Harsh Satija, Joshua Romoff +2

Randomized value functions offer a promising approach towards the challenge of efficient exploration in complex environments with high dimensional state and action spaces. Unlike t…

cs.LG2018

Decoupling Dynamics and Reward for Transfer Learning

Amy Zhang, Harsh Satija, Joelle Pineau

Current reinforcement learning (RL) methods can successfully learn single tasks but often generalize poorly to modest perturbations in task domain or training procedure. In this wo…