1 paper
Haque Ishfaq, Guangyuan Wang, Sami Nur Islam +1
Existing actor-critic algorithms, which are popular for continuous control reinforcement learning (RL) tasks, suffer from poor sample efficiency due to lack of principled explorati…