activity
20212024
most citedUnsupervised Reinforcement Learning in Multiple Environments

3 citations · 4 across the 6 of their papers we have counts for

collaborators

6 papers

cs.LG2024

Geometric Active Exploration in Markov Decision Processes: the Benefit of Abstraction

Riccardo De Santi, Federico Arangath Joseph, Noah Liniger +2

How can a scientist use a Reinforcement Learning (RL) algorithm to design experiments over a dynamical system's state space? In the case of finite and Markovian systems, an area ca…

cs.LG20241 cited

The Limits of Pure Exploration in POMDPs: When the Observation Entropy is Enough

Riccardo Zamboni, Duilio Cirino, Marcello Restelli +1

The problem of pure exploration in Markov decision processes has been cast as maximizing the entropy over the state distribution induced by the agent's policy, an objective that ha…

cs.LG2024

How to Explore with Belief: State Entropy Maximization in POMDPs

Riccardo Zamboni, Duilio Cirino, Marcello Restelli +1

Recent works have studied *state entropy maximization* in reinforcement learning, in which the agent's objective is to learn a policy inducing high entropy over states visitation (…

cs.LG2024

Test-Time Regret Minimization in Meta Reinforcement Learning

Mirco Mutti, Aviv Tamar

Meta reinforcement learning sets a distribution over a set of tasks on which the agent can train at will, then is asked to learn an optimal policy for any test task efficiently. In…

cs.LG2023

A Tale of Sampling and Estimation in Discounted Reinforcement Learning

Alberto Maria Metelli, Mirco Mutti, Marcello Restelli

The most relevant problems in discounted reinforcement learning involve estimating the mean of a function under the stationary distribution of a Markov reward process, such as the…

cs.LG20213 cited

Unsupervised Reinforcement Learning in Multiple Environments

Mirco Mutti, Mattia Mancassola, Marcello Restelli

Several recent works have been dedicated to unsupervised reinforcement learning in a single environment, in which a policy is first pre-trained with unsupervised interactions, and…