1 paper · 1 filter
Brent Yi, Hongsuk Choi, Himanshu Gaurav Singh +9
Likelihood-based policy gradient methods are the dominant approach for training robot control policies from rewards. These methods rely on differentiable action likelihoods, which…