1 paper
Taisuke Kobayashi, Takumi Aotani
This paper proposes a new design method for a stochastic control policy using a normalizing flow (NF). In reinforcement learning (RL), the policy is usually modeled as a distributi…