5 papers
Distributions as Actions: A Unified Framework for Diverse Action Spaces
Jiamin He, A. Rupam Mahmood, Martha White
We introduce a novel reinforcement learning (RL) framework that treats parameterized action distributions as actions, redefining the boundary between agent and environment. This re…
Revisiting Mixture Policies in Entropy-Regularized Actor-Critic
Jiamin He, Samuel Neumann, Jincheng Mei +2
Mixture policies theoretically offer greater flexibility than unimodal policies in continuous action reinforcement learning, but the practical benefits of this complexity remain el…
Extending Differential Temporal Difference Methods for Episodic Problems
Kris De Asis, Mohamed Elsayed, Jiamin He
Differential temporal difference (TD) methods are value-based reinforcement learning algorithms that have been proposed for infinite-horizon problems. They rely on reward centering…
Forager: a lightweight testbed for continual learning with partial observability in RL
Steven Tang, Xinze Xiong, Anna Hakhverdyan +7
In continual reinforcement learning (CRL), good performance requires never-ending learning, acting, and exploration in a big, partially observable world. Most CRL experiments have…
Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers
Gautham Vasan, Mohamed Elsayed, Alireza Azimi +5
Modern deep policy gradient methods achieve effective performance on simulated robotic tasks, but they all require large replay buffers or expensive batch updates, or both, making…