1 paper
Zeyuan Wang, Da Li, Yulin Chen +6
Online off-policy reinforcement learning (RL) is shaped by two coupled choices: the policy class and the update rule. Gaussian policies are fast and have tractable entropy, but str…