1 paper
Zheng Tian, Ying Wen, Zhichen Gong +3
In a single-agent setting, reinforcement learning (RL) tasks can be cast into an inference problem by introducing a binary random variable o, which stands for the "optimality". In…