activity
20162020
most citedConstrained Policy Optimization

112 citations · 315 across the 4 of their papers we have counts for

collaborators

7 papers

math.OC202044 cited

Responsive Safety in Reinforcement Learning by PID Lagrangian Methods

Adam Stooke, Joshua Achiam, Pieter Abbeel

Lagrangian methods are widely used algorithms for constrained optimization problems, but their learning dynamics exhibit oscillations and overshoot which, when applied to safe rein…

cs.LG201960 cited

Towards Characterizing Divergence in Deep Q-Learning

Joshua Achiam, Ethan Knight, Pieter Abbeel

Deep Q-Learning (DQL), a family of temporal difference algorithms for control, employs three techniques collectively known as the `deadly triad' in reinforcement learning: bootstra…

cs.AI2018

Variational Option Discovery Algorithms

Joshua Achiam, Harrison Edwards, Dario Amodei +1

We explore methods for option discovery based on variational inference and make two algorithmic contributions. First: we highlight a tight connection between variational option dis…

cs.LG2018

On First-Order Meta-Learning Algorithms

Alex Nichol, Joshua Achiam, John Schulman

This paper considers meta-learning problems, where there is a distribution of tasks, and we would like to obtain an agent that performs well (i.e., learns quickly) when presented w…

cs.LG2017112 cited

Constrained Policy Optimization

Joshua Achiam, David Held, Aviv Tamar +1

For many applications of reinforcement learning it can be more convenient to specify both a reward function and constraints, rather than trying to design behavior through the rewar…

cs.LG201799 cited

Surprise-Based Intrinsic Motivation for Deep Reinforcement Learning

Joshua Achiam, Shankar Sastry

Exploration in complex domains is a key challenge in reinforcement learning, especially for tasks with very sparse rewards. Recent successes in deep reinforcement learning have bee…