1 paper
Yiyang He, Zhichun Zhou, Ziwei Wang +2
Many continuous-control policies are optimized as unbounded Gaussians and then mapped into bounded actions. We show that where entropy is measured changes the policy geometry learn…