1 paper
Rijul Tandon, Peter Vamplew, Cameron Foale
In most value-based reinforcement learning (RL) algorithms, the agent estimates only the expected reward for each action and selects the action with the highest reward. In contrast…