1 citations · 1 across the 1 of their papers we have counts for
1 paper
Haonan Yu, Haichao Zhang, Wei Xu
Maximum entropy (MaxEnt) RL maximizes a combination of the original task reward and an entropy reward. It is believed that the regularization imposed by entropy, on both policy imp…