112 citations · 315 across the 4 of their papers we have counts for
7 papers
Responsive Safety in Reinforcement Learning by PID Lagrangian Methods
Adam Stooke, Joshua Achiam, Pieter Abbeel
Lagrangian methods are widely used algorithms for constrained optimization problems, but their learning dynamics exhibit oscillations and overshoot which, when applied to safe rein…
Towards Characterizing Divergence in Deep Q-Learning
Joshua Achiam, Ethan Knight, Pieter Abbeel
Deep Q-Learning (DQL), a family of temporal difference algorithms for control, employs three techniques collectively known as the `deadly triad' in reinforcement learning: bootstra…
Variational Option Discovery Algorithms
Joshua Achiam, Harrison Edwards, Dario Amodei +1
We explore methods for option discovery based on variational inference and make two algorithmic contributions. First: we highlight a tight connection between variational option dis…
On First-Order Meta-Learning Algorithms
Alex Nichol, Joshua Achiam, John Schulman
This paper considers meta-learning problems, where there is a distribution of tasks, and we would like to obtain an agent that performs well (i.e., learns quickly) when presented w…
Constrained Policy Optimization
Joshua Achiam, David Held, Aviv Tamar +1
For many applications of reinforcement learning it can be more convenient to specify both a reward function and constraints, rather than trying to design behavior through the rewar…
Surprise-Based Intrinsic Motivation for Deep Reinforcement Learning
Joshua Achiam, Shankar Sastry
Exploration in complex domains is a key challenge in reinforcement learning, especially for tasks with very sparse rewards. Recent successes in deep reinforcement learning have bee…