1 paper
Jincheng Mei, Ian Osband
Softmax policy gradient converges at O(1/t), but its transient behavior near sub-optimal corners of the simplex can be exponentially slow. The bottleneck is self-trapping: negati…