301 citations · 358 across the 12 of their papers we have counts for
1 paper · 2 filters
Silviu Pitis
The softmax policy π(a∣s)∝exp(βQ(s,a)) is the default model of stochastic choice in reinforcement learning (RL). Various justifications based on robustness, explora…