158 citations
1 paper
Freek Stulp, Olivier Sigaud
There has been a recent focus in reinforcement learning on addressing continuous state and action problems by optimizing parameterized policies. PI2 is a recent example of this app…