1 paper
Seiji Ishihara, Harukazu Igarashi
A method of a fusion of fuzzy inference and policy gradient reinforcement learning has been proposed that directly learns, as maximizes the expected value of the reward per episode…