4 papers
Extending Differential Temporal Difference Methods for Episodic Problems
Kris De Asis, Mohamed Elsayed, Jiamin He
Differential temporal difference (TD) methods are value-based reinforcement learning algorithms that have been proposed for infinite-horizon problems. They rely on reward centering…
Intentional Updates for Streaming Reinforcement Learning
Arsalan Sharifnassab, Mohamed Elsayed, Kris De Asis +2
In gradient-based learning, a step size chosen in parameter units does not produce a predictable per-step change in function output. This often leads to instability in the streamin…
Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers
Gautham Vasan, Mohamed Elsayed, Alireza Azimi +5
Modern deep policy gradient methods achieve effective performance on simulated robotic tasks, but they all require large replay buffers or expensive batch updates, or both, making…
Streaming Deep Reinforcement Learning Finally Works
Mohamed Elsayed, Gautham Vasan, A. Rupam Mahmood
Natural intelligence processes experience as a continuous stream, sensing, acting, and learning moment-by-moment in real time. Streaming learning, the modus operandi of classic rei…