1 paper
Yao Lyu, Xiangteng Zhang, Shengbo Eben Li +5
Training deep reinforcement learning (RL) agents necessitates overcoming the highly unstable nonconvex stochastic optimization inherent in the trial-and-error mechanism. To tackle…