activity
20212023
most citedTruncating Trajectories in Monte Carlo Reinforcement Learning

3 citations · 8 across the 10 of their papers we have counts for

collaborators

10 papers

cs.LG2023

An Option-Dependent Analysis of Regret Minimization Algorithms in Finite-Horizon Semi-Markov Decision Processes

Gianluca Drappo, Alberto Maria Metelli, Marcello Restelli

A large variety of real-world Reinforcement Learning (RL) tasks is characterized by a complex and heterogeneous structure that makes end-to-end (or flat) approaches hardly applicab…

cs.LG20233 cited

Truncating Trajectories in Monte Carlo Reinforcement Learning

Riccardo Poiani, Alberto Maria Metelli, Marcello Restelli

In Reinforcement Learning (RL), an agent acts in an unknown environment to maximize the expected cumulative discounted sum of an external reward signal, i.e., the expected return.…

cs.LG20233 cited

Towards Theoretical Understanding of Inverse Reinforcement Learning

Alberto Maria Metelli, Filippo Lazzati, Marcello Restelli

Inverse reinforcement learning (IRL) denotes a powerful family of algorithms for recovering a reward function justifying the behavior demonstrated by an expert agent. A well-known…

cs.LG2023

A Tale of Sampling and Estimation in Discounted Reinforcement Learning

Alberto Maria Metelli, Mirco Mutti, Marcello Restelli

The most relevant problems in discounted reinforcement learning involve estimating the mean of a function under the stationary distribution of a Markov reward process, such as the…

cs.LG2023

Interpretable Linear Dimensionality Reduction based on Bias-Variance Analysis

Paolo Bonetti, Alberto Maria Metelli, Marcello Restelli

One of the central issues of several machine learning applications on real data is the choice of the input features. Ideally, the designer should select only the relevant, non-redu…

cs.LG2023

Information-Theoretic Regret Bounds for Bandits with Fixed Expert Advice

Khaled Eldowa, Nicolò Cesa-Bianchi, Alberto Maria Metelli +1

We investigate the problem of bandits with expert advice when the experts are fixed and known distributions over the actions. Improving on previous analyses, we show that the regre…