6 citations · 6 across the 3 of their papers we have counts for
1 paper · 1 filter
Long Yang, Zhixiong Huang, Fenghao Lei +6
Popular reinforcement learning (RL) algorithms tend to produce a unimodal policy distribution, which weakens the expressiveness of complicated policy and decays the ability of expl…