2 papers
cs.LG2020
Task-Agnostic Exploration via Policy Gradient of a Non-Parametric State Entropy Estimate
Mirco Mutti, Lorenzo Pratissoli, Marcello Restelli
In a reward-free environment, what is a suitable intrinsic objective for an agent to pursue so that it can learn an optimal task-agnostic exploration policy? In this paper, we argu…
cs.LG2019
An Intrinsically-Motivated Approach for Learning Highly Exploring and Fast Mixing Policies
Mirco Mutti, Marcello Restelli
What is a good exploration strategy for an agent that interacts with an environment in the absence of external rewards? Ideally, we would like to get a policy driving towards a uni…