1 paper
Erfan Entezami, Mahsa Sahebdel, Dhawal Gupta
Training a model-free reinforcement learning agent requires allowing the agent to sufficiently explore the environment to search for an optimal policy. In safety-constrained enviro…