1 paper
Kai Jiang, XiaoLong Qin
Reinforcement learning usually uses the feedback rewards of environmental to train agents. But the rewards in the actual environment are sparse, and even some environments will not…