9 citations · 13 across the 3 of their papers we have counts for
4 papers
A unified view of likelihood ratio and reparameterization gradients
Paavo Parmas, Masashi Sugiyama
Reparameterization (RP) and likelihood ratio (LR) gradient estimators are used to estimate gradients of expectations throughout machine learning and reinforcement learning; however…
A unified view of likelihood ratio and reparameterization gradients and an optimal importance sampling scheme
Paavo Parmas, Masashi Sugiyama
Reparameterization (RP) and likelihood ratio (LR) gradient estimators are used throughout machine and reinforcement learning; however, they are usually explained as simple mathemat…
Total stochastic gradient algorithms and applications in reinforcement learning
Paavo Parmas
Backpropagation and the chain rule of derivatives have been prominent; however, the total derivative rule has not enjoyed the same amount of attention. In this work we show how the…
PIPPS: Flexible Model-Based Policy Search Robust to the Curse of Chaos
Paavo Parmas, Carl Edward Rasmussen, Jan Peters +1
Previously, the exploding gradient problem has been explained to be central in deep learning and model-based reinforcement learning, because it causes numerical issues and instabil…