1 paper · 1 filter
Ruixuan Miao, Xu Lu, Cong Tian +2
The commonly used Reinforcement Learning (RL) model, MDPs (Markov Decision Processes), has a basic premise that rewards depend on the current state and action only. However, many r…