1 paper · 1 filter
Daoming Lyu, Fangkai Yang, Bo Liu +1
Conventional reinforcement learning (RL) allows an agent to learn policies via environmental rewards only, with a long and slow learning curve, especially at the beginning stage. O…