1 paper · 1 filter
Gal Dalal, Assaf Hallak, Gugan Thoppe +2
Policy gradient methods are notorious for having a large variance and high sample complexity. To mitigate this, we introduce SoftTreeMax -- a generalization of softmax that employs…