1 paper
Ruixuan Miao, Xu Lu, Cong Tian +2
The commonly used Reinforcement Learning (RL) model, MDPs (Markov Decision Processes), has a basic premise that rewards depend on the current state and action only. However, many r…