1 paper
Konatsu Miyamoto, Masaya Suzuki, Yuma Kigami +1
In this paper, as a study of reinforcement learning, we converge the Q function to unbounded rewards such as Gaussian distribution. From the central limit theorem, in some real-wor…