157 citations · 546 across the 8 of their papers we have counts for
1 paper · 1 filter
Zhao Song, Ronald E. Parr, Lawrence Carin
The impact of softmax on the value function itself in reinforcement learning (RL) is often viewed as problematic because it leads to sub-optimal value (or Q) functions and interfer…