papers

Publications (103)

cs.LG2021

Unsupervised Reinforcement Learning in Multiple Environments

Mirco Mutti, Mattia Mancassola, Marcello Restelli

Several recent works have been dedicated to unsupervised reinforcement learning in a single environment, in which a policy is first pre-trained with unsupervised interactions, and…

cs.LG2021

Task-Agnostic Exploration via Policy Gradient of a Non-Parametric State Entropy Estimate

Mirco Mutti, Lorenzo Pratissoli, Marcello Restelli

In a reward-free environment, what is a suitable intrinsic objective for an agent to pursue so that it can learn an optimal task-agnostic exploration policy? In this paper, we argu…

cs.LG2020

Control Frequency Adaptation via Action Persistence in Batch Reinforcement Learning

Alberto Maria Metelli, Flavio Mazzolini, Lorenzo Bisi +2

The choice of the control frequency of a system has a relevant impact on the ability of reinforcement learning algorithms to learn a highly performing policy. In this paper, we int…

cs.LG2018

Importance Weighted Transfer of Samples in Reinforcement Learning

Andrea Tirinzoni, Andrea Sessa, Matteo Pirotta +1

We consider the transfer of experience samples (i.e., tuples < s, a, s', r >) in reinforcement learning (RL), collected from a set of source tasks to improve the learning process i…

cs.LG2022

Storehouse: a Reinforcement Learning Environment for Optimizing Warehouse Management

Julen Cestero, Marco Quartulli, Alberto Maria Metelli +1

Warehouse Management Systems have been evolving and improving thanks to new Data Intelligence techniques. However, many current optimizations have been applied to specific cases or…

q-fin.TR2024

Exploiting Risk-Aversion and Size-dependent fees in FX Trading with Fitted Natural Actor-Critic

Vito Alessandro Monaco, Antonio Riva, Luca Sabbioni +5

In recent years, the popularity of artificial intelligence has surged due to its widespread application in various fields. The financial sector has harnessed its advantages for mul…

cs.LG2021

Meta-Reinforcement Learning by Tracking Task Non-stationarity

Riccardo Poiani, Andrea Tirinzoni, Marcello Restelli

Many real-world domains are subject to a structured non-stationarity which affects the agent's goals and the environmental dynamics. Meta-reinforcement learning (RL) has been shown…

eess.SY2024

State and Action Factorization in Power Grids

Gianvito Losapio, Davide Beretta, Marco Mussi +2

The increase of renewable energy generation towards the zero-emission target is making the problem of controlling power grids more and more challenging. The recent series of compet…

cs.LG2018

Stochastic Variance-Reduced Policy Gradient

Matteo Papini, Damiano Binaghi, Giuseppe Canonaco +2

In this paper, we propose a novel reinforcement- learning algorithm consisting in a stochastic variance-reduced version of policy gradient for solving Markov Decision Processes (MD…

cs.LG2021

Newton Optimization on Helmholtz Decomposition for Continuous Games

Giorgia Ramponi, Marcello Restelli

Many learning problems involve multiple agents optimizing different interactive functions. In these problems, the standard policy gradient algorithms fail due to the non-stationari…

cs.LG2020

Policy Optimization as Online Learning with Mediator Feedback

Alberto Maria Metelli, Matteo Papini, Pierluca D'Oro +1

Policy Optimization (PO) is a widely used approach to address continuous control tasks. In this paper, we introduce the notion of mediator feedback that frames PO as an online lear…

cs.LG2025

Achieving Regret in Average-Reward POMDPs with Known Observation Models

Alessio Russo, Alberto Maria Metelli, Marcello Restelli

We tackle average-reward infinite-horizon POMDPs with an unknown transition model but a known observation model, a setting that has been previously addressed in two limiting ways:…

cs.LG2024

Exploiting Causal Graph Priors with Posterior Sampling for Reinforcement Learning

Mirco Mutti, Riccardo De Santi, Marcello Restelli +2

Posterior sampling allows exploitation of prior knowledge on the environment's transition dynamics to improve the sample efficiency of reinforcement learning. The prior is typicall…

cs.LG2023

Stepsize Learning for Policy Gradient Methods in Contextual Markov Decision Processes

Luca Sabbioni, Francesco Corda, Marcello Restelli

Policy-based algorithms are among the most widely adopted techniques in model-free RL, thanks to their strong theoretical groundings and good properties in continuous action spaces…

cs.LG2024

The Limits of Pure Exploration in POMDPs: When the Observation Entropy is Enough

Riccardo Zamboni, Duilio Cirino, Marcello Restelli +1

The problem of pure exploration in Markov decision processes has been cast as maximizing the entropy over the state distribution induced by the agent's policy, an objective that ha…

cs.LG2022

Dynamic Pricing with Volume Discounts in Online Settings

Marco Mussi, Gianmarco Genalti, Alessandro Nuara +3

According to the main international reports, more pervasive industrial and business-process automation, thanks to machine learning and advanced analytic tools, will unlock more tha…

stat.ML2026

Bridging Rested and Restless Bandits with Graph-Triggering: Rising and Rotting

Gianmarco Genalti, Marco Mussi, Nicola Gatti +3

Rested and Restless Bandits are two well-known bandit settings that are useful to model real-world sequential decision-making problems in which the expected reward of an arm evolve…

cs.LG2024

Policy Gradient with Active Importance Sampling

Matteo Papini, Giorgio Manganini, Alberto Maria Metelli +1

Importance sampling (IS) represents a fundamental technique for a large surge of off-policy reinforcement learning approaches. Policy gradient (PG) methods, in particular, signific…

cs.LG2025

Towards Principled Unsupervised Multi-Agent Reinforcement Learning

Riccardo Zamboni, Mirco Mutti, Marcello Restelli

In reinforcement learning, we typically refer to unsupervised pre-training when we aim to pre-train a policy without a priori access to the task specification, i.e. rewards, to be…

cs.LG2024

Local Linearity: the Key for No-regret Reinforcement Learning in Continuous MDPs

Davide Maran, Alberto Maria Metelli, Matteo Papini +1

Achieving the no-regret property for Reinforcement Learning (RL) problems in continuous state and action-space environments is one of the major open problems in the field. Existing…

cs.LG2024

Projection by Convolution: Optimal Sample Complexity for Reinforcement Learning in Continuous-Space MDPs

Davide Maran, Alberto Maria Metelli, Matteo Papini +1

We consider the problem of learning an -optimal policy in a general class of continuous-space Markov decision processes (MDPs) having smooth Bellman operators. Given a…

cs.LG2024

Pure Exploration under Mediators' Feedback

Riccardo Poiani, Alberto Maria Metelli, Marcello Restelli

Stochastic multi-armed bandits are a sequential-decision-making framework, where, at each interaction step, the learner selects an arm and observes a stochastic reward. Within the…

cs.LG2025

Online Dynamic Pricing of Complementary Products

Marco Mussi, Marcello Restelli

Traditional pricing paradigms, once dominated by static models and rule-based heuristics, are increasingly being replaced by dynamic, data-driven approaches powered by machine lear…

cs.LG2024

Efficient Learning of POMDPs with Known Observation Model in Average-Reward Setting

Alessio Russo, Alberto Maria Metelli, Marcello Restelli

Dealing with Partially Observable Markov Decision Processes is notably a challenging task. We face an average-reward infinite-horizon POMDP setting with an unknown transition model…

cs.LG2022

Delayed Reinforcement Learning by Imitation

Pierre Liotet, Davide Maran, Lorenzo Bisi +1

When the agent's observations or interactions are delayed, classic reinforcement learning tools usually fail. In this paper, we propose a simple yet new and efficient solution to t…

cs.LG2021

Leveraging Good Representations in Linear Contextual Bandits

Matteo Papini, Andrea Tirinzoni, Marcello Restelli +2

The linear contextual bandit literature is mostly focused on the design of efficient learning algorithms for a given representation. However, a contextual bandit problem may admit…

cs.LG2024

Sharing Knowledge in Multi-Task Deep Reinforcement Learning

Carlo D'Eramo, Davide Tateo, Andrea Bonarini +2

We study the benefit of sharing representations among tasks to enable the effective use of deep neural networks in Multi-Task Reinforcement Learning. We leverage the assumption tha…

cs.LG2024

Autoregressive Bandits

Francesco Bacchiocchi, Gianmarco Genalti, Davide Maran +4

Autoregressive processes naturally arise in a large variety of real-world scenarios, including stock markets, sales forecasting, weather prediction, advertising, and pricing. When…

cs.LG2021

Lifelong Hyper-Policy Optimization with Multiple Importance Sampling Regularization

Pierre Liotet, Francesco Vidaich, Alberto Maria Metelli +1

Learning in a lifelong setting, where the dynamics continually evolve, is a hard challenge for current reinforcement learning algorithms. Yet this would be a much needed feature fo…

cs.LG2019

Gradient-Aware Model-based Policy Search

Pierluca D'Oro, Alberto Maria Metelli, Andrea Tirinzoni +2

Traditional model-based reinforcement learning approaches learn a model of the environment dynamics without explicitly considering how it will be used by the agent. In the presence…

cs.LG2020

MushroomRL: Simplifying Reinforcement Learning Research

Carlo D'Eramo, Davide Tateo, Andrea Bonarini +2

MushroomRL is an open-source Python library developed to simplify the process of implementing and running Reinforcement Learning (RL) experiments. Compared to other available libra…

cs.LG2019

Policy Space Identification in Configurable Environments

Alberto Maria Metelli, Guglielmo Manneschi, Marcello Restelli

We study the problem of identifying the policy space of a learning agent, having access to a set of demonstrations generated by its optimal policy. We introduce an approach based o…

cs.RO2024

A Retrospective on the Robot Air Hockey Challenge: Benchmarking Robust, Reliable, and Safe Learning Techniques for Real-world Robotics

Puze Liu, Jonas Günster, Niklas Funk +17

Machine learning methods have a groundbreaking impact in many application domains, but their application on real robotic platforms is still limited. Despite the many challenges ass…

cs.LG2019

Feature Selection via Mutual Information: New Theoretical Insights

Mario Beraha, Alberto Maria Metelli, Matteo Papini +2

Mutual information has been successfully adopted in filter feature-selection methods to assess both the relevancy of a subset of features in predicting the target variable and the…

cs.LG2020

Sequential Transfer in Reinforcement Learning with a Generative Model

Andrea Tirinzoni, Riccardo Poiani, Marcello Restelli

We are interested in how to design reinforcement learning agents that provably reduce the sample complexity for learning new tasks by transferring knowledge from previously-solved…

cs.LG2020

Time-Variant Variational Transfer for Value Functions

Giuseppe Canonaco, Andrea Soprani, Manuel Roveri +1

In most of the transfer learning approaches to reinforcement learning (RL) the distribution over the tasks is assumed to be stationary. Therefore, the target and source tasks are i…

cs.LG2025

Building surrogate models using trajectories of agents trained by Reinforcement Learning

Julen Cestero, Marco Quartulli, Marcello Restelli

Sample efficiency in the face of computationally expensive simulations is a common concern in surrogate modeling. Current strategies to minimize the number of samples needed are no…

cs.LG2026

Unsupervised Behavioral Compression: Learning Low-Dimensional Policy Manifolds through State-Occupancy Matching

Andrea Fraschini, Davide Tenedini, Riccardo Zamboni +2

Deep Reinforcement Learning (DRL) is widely recognized as sample-inefficient, a limitation attributable in part to the high dimensionality and substantial functional redundancy inh…

cs.LG2023

Causal Feature Selection via Transfer Entropy

Paolo Bonetti, Alberto Maria Metelli, Marcello Restelli

Machine learning algorithms are designed to capture complex relationships between features. In this context, the high dimensionality of data often results in poor model performance…

cs.LG2023

A Tale of Sampling and Estimation in Discounted Reinforcement Learning

Alberto Maria Metelli, Mirco Mutti, Marcello Restelli

The most relevant problems in discounted reinforcement learning involve estimating the mean of a function under the stationary distribution of a Markov reward process, such as the…

cs.LG2022

ARLO: A Framework for Automated Reinforcement Learning

Marco Mussi, Davide Lombarda, Alberto Maria Metelli +2

Automated Reinforcement Learning (AutoRL) is a relatively new area of research that is gaining increasing attention. The objective of AutoRL consists in easing the employment of Re…

cs.LG2026

K-Myriad: Jump-starting reinforcement learning with unsupervised parallel agents

Vincenzo De Paola, Mirco Mutti, Riccardo Zamboni +1

Parallelization in Reinforcement Learning is typically employed to speed up the training of a single policy, where multiple workers collect experience from an identical sampling di…

cs.LG2025

Limitations of Physics-Informed Neural Networks: a Study on Smart Grid Surrogation

Julen Cestero, Carmine Delle Femine, Kenji S. Muro +2

Physics-Informed Neural Networks (PINNs) present a transformative approach for smart grid modeling by integrating physical laws directly into learning frameworks, addressing critic…

cs.LG2024

Parameterized Projected Bellman Operator

Théo Vincent, Alberto Maria Metelli, Boris Belousov +3

Approximate value iteration (AVI) is a family of algorithms for reinforcement learning (RL) that aims to obtain an approximation of the optimal value function. Generally, AVI algor…

cs.LG2024

Information Capacity Regret Bounds for Bandits with Mediator Feedback

Khaled Eldowa, Nicolò Cesa-Bianchi, Alberto Maria Metelli +1

This work addresses the mediator feedback problem, a bandit game where the decision set consists of a number of policies, each associated with a probability distribution over a com…

cs.LG2019

An Intrinsically-Motivated Approach for Learning Highly Exploring and Fast Mixing Policies

Mirco Mutti, Marcello Restelli

What is a good exploration strategy for an agent that interacts with an environment in the absence of external rewards? Ideally, we would like to get a policy driving towards a uni…

cs.LG2023

Information-Theoretic Regret Bounds for Bandits with Fixed Expert Advice

Khaled Eldowa, Nicolò Cesa-Bianchi, Alberto Maria Metelli +1

We investigate the problem of bandits with expert advice when the experts are fixed and known distributions over the actions. Improving on previous analyses, we show that the regre…

cs.LG2024

How to Explore with Belief: State Entropy Maximization in POMDPs

Riccardo Zamboni, Duilio Cirino, Marcello Restelli +1

Recent works have studied *state entropy maximization* in reinforcement learning, in which the agent's objective is to learn a policy inducing high entropy over states visitation (…

cs.LG2024

Best Arm Identification for Stochastic Rising Bandits

Marco Mussi, Alessandro Montenegro, Francesco Trovó +2

Stochastic Rising Bandits (SRBs) model sequential decision-making problems in which the expected reward of the available options increases every time they are selected. This settin…

cs.LG2026

Online Market Making and the Value of Observing the Order Book

Davide Maran, Marcello Restelli

We study an online market-making problem in which a learner sequentially posts bid and ask prices for a single asset while interacting with traders holding private valuations. Unli…

cs.LG2025

Optimal Multi-Fidelity Best-Arm Identification

Riccardo Poiani, Rémy Degenne, Emilie Kaufmann +2

In bandit best-arm identification, an algorithm is tasked with finding the arm with highest mean reward with a specified accuracy as fast as possible. We study multi-fidelity best-…

cs.LG2023

Interpretable Linear Dimensionality Reduction based on Bias-Variance Analysis

Paolo Bonetti, Alberto Maria Metelli, Marcello Restelli

One of the central issues of several machine learning applications on real data is the choice of the input features. Ideally, the designer should select only the relevant, non-redu…

cs.LG2020

Inverse Reinforcement Learning from a Gradient-based Learner

Giorgia Ramponi, Gianluca Drappo, Marcello Restelli

Inverse Reinforcement Learning addresses the problem of inferring an expert's reward function from demonstrations. However, in many applications, we not only have access to the exp…

cs.LG2022

Optimizing Empty Container Repositioning and Fleet Deployment via Configurable Semi-POMDPs

Riccardo Poiani, Ciprian Stirbu, Alberto Maria Metelli +1

With the continuous growth of the global economy and markets, resource imbalance has risen to be one of the central issues in real logistic scenarios. In marine transportation, thi…

cs.AI2011

Transfer from Multiple MDPs

Alessandro Lazaric, Marcello Restelli

Transfer reinforcement learning (RL) methods leverage on the experience collected on a set of source tasks to speed-up RL algorithms. A simple and effective approach is to transfer…

cs.LG2022

Simultaneously Updating All Persistence Values in Reinforcement Learning

Luca Sabbioni, Luca Al Daire, Lorenzo Bisi +2

In reinforcement learning, the performance of learning agents is highly sensitive to the choice of time discretization. Agents acting at high frequencies have the best control oppo…

cs.LG2022

Tight Performance Guarantees of Imitator Policies with Continuous Actions

Davide Maran, Alberto Maria Metelli, Marcello Restelli

Behavioral Cloning (BC) aims at learning a policy that mimics the behavior demonstrated by an expert. The current theoretical understanding of BC is limited to the case of finite a…

cs.AI2018

Configurable Markov Decision Processes

Alberto Maria Metelli, Mirco Mutti, Marcello Restelli

In many real-world problems, there is the possibility to configure, to a limited extent, some environmental parameters to improve the performance of a learning agent. In this paper…

cs.LG2025

Power Grid Control with Graph-Based Distributed Reinforcement Learning

Carlo Fabrizio, Gianvito Losapio, Marco Mussi +2

The necessary integration of renewable energy sources, combined with the expanding scale of power networks, presents significant challenges in controlling modern power grids. Tradi…

cs.LG2022

Analysis, Characterization, Prediction and Attribution of Extreme Atmospheric Events with Machine Learning: a Review

Sancho Salcedo-Sanz, Jorge Pérez-Aracil, Guido Ascenso +9

Atmospheric Extreme Events (EEs) cause severe damages to human societies and ecosystems. The frequency and intensity of EEs and other associated events are increasing in the curren…

cs.LG2020

A Novel Confidence-Based Algorithm for Structured Bandits

Andrea Tirinzoni, Alessandro Lazaric, Marcello Restelli

We study finite-armed stochastic bandits where the rewards of each arm might be correlated to those of other arms. We introduce a novel phased algorithm that exploits the given str…

quant-ph2021

Quantum Compiling by Deep Reinforcement Learning

Lorenzo Moro, Matteo G. A. Paris, Marcello Restelli +1

The architecture of circuital quantum computers requires computing layers devoted to compiling high-level quantum algorithms into lower-level circuits of quantum gates. The general…

cs.LG2024

Truncating Trajectories in Monte Carlo Policy Evaluation: an Adaptive Approach

Riccardo Poiani, Nicole Nobili, Alberto Maria Metelli +1

Policy evaluation via Monte Carlo (MC) simulation is at the core of many MC Reinforcement Learning (RL) algorithms (e.g., policy gradient methods). In this context, the designer of…

cs.LG2023

Provably Efficient Causal Model-Based Reinforcement Learning for Systematic Generalization

Mirco Mutti, Riccardo De Santi, Emanuele Rossi +3

In the sequential decision making setting, an agent aims to achieve systematic generalization over a large, possibly infinite, set of environments. Such environments are modeled as…

cs.LG2026

Finite Sample Bounds for Non-Parametric Regression: Optimal Sample Efficiency and Space Complexity

Davide Maran, Marcello Restelli

We address the problem of learning an unknown smooth function and its derivatives from noisy pointwise evaluations under the supremum norm. While classical nonparametric regression…

cs.LG2023

Dynamical Linear Bandits

Marco Mussi, Alberto Maria Metelli, Marcello Restelli

In many real-world sequential decision-making problems, an action does not immediately reflect on the feedback and spreads its effects over a long time frame. For instance, in onli…

cs.LG2022

Stochastic Rising Bandits

Alberto Maria Metelli, Francesco Trovò, Matteo Pirola +1

This paper is in the field of stochastic Multi-Armed Bandits (MABs), i.e., those sequential selection techniques able to learn online using only the feedback given by the chosen op…

cs.LG2019

Risk-Averse Trust Region Optimization for Reward-Volatility Reduction

Lorenzo Bisi, Luca Sabbioni, Edoardo Vittori +2

In real-world decision-making problems, for instance in the fields of finance, robotics or autonomous driving, keeping uncertainty under control is as important as maximizing expec…

cs.LG2025

Gym4ReaL: A Suite for Benchmarking Real-World Reinforcement Learning

Davide Salaorni, Vincenzo De Paola, Samuele Delpero +9

In recent years, \emph{Reinforcement Learning} (RL) has made remarkable progress, achieving superhuman performance in a wide range of simulated environments. As research moves towa…

cs.LG2025

Scalable Multi-Agent Offline Reinforcement Learning and the Role of Information

Riccardo Zamboni, Enrico Brunetti, Marcello Restelli

Offline Reinforcement Learning (RL) focuses on learning policies solely from a batch of previously collected data. offering the potential to leverage such datasets effectively with…

cs.LG2021

Reinforcement Learning in Linear MDPs: Constant Regret and Representation Selection

Matteo Papini, Andrea Tirinzoni, Aldo Pacchiano +3

We study the role of the representation of state-action value functions in regret minimization in finite-horizon Markov Decision Processes (MDPs) with linear structure. We first de…

cs.LG2025

A Provably Efficient Option-Based Algorithm for both High-Level and Low-Level Learning

Gianluca Drappo, Alberto Maria Metelli, Marcello Restelli

Hierarchical Reinforcement Learning (HRL) approaches have shown successful results in solving a large variety of complex, structured, long-horizon problems. Nevertheless, a full th…

cs.LG2017

Cost-Sensitive Approach to Batch Size Adaptation for Gradient Descent

Matteo Pirotta, Marcello Restelli

In this paper, we propose a novel approach to automatically determine the batch size in stochastic gradient descent methods. The choice of the batch size induces a trade-off betwee…

cs.LG2025

A Reinforcement Learning Approach for Optimal Control in Microgrids

Davide Salaorni, Federico Bianchi, Francesco Trovò +1

The increasing integration of renewable energy sources (RESs) is transforming traditional power grid networks, which require new approaches for managing decentralized energy produc…

quant-ph2019

Coherent Transport of Quantum States by Deep Reinforcement Learning

Riccardo Porotti, Dario Tamascelli, Marcello Restelli +1

Some problems in physics can be handled only after a suitable \textit{ansatz }solution has been guessed. Such method is therefore resilient to generalization, resulting of limited…

cs.AI2014

Multi-objective Reinforcement Learning with Continuous Pareto Frontier Approximation Supplementary Material

Matteo Pirotta, Simone Parisi, Marcello Restelli

This document contains supplementary material for the paper "Multi-objective Reinforcement Learning with Continuous Pareto Frontier Approximation", published at the Twenty-Ninth AA…

cs.GT2013

Efficient evolutionary dynamics with extensive-form games

Nicola Gatti, Fabio Panozzo, Marcello Restelli

Evolutionary game theory combines game theory and dynamical systems and is customarily adopted to describe evolutionary dynamics in multi-agent systems. In particular, it has been…

cs.LG2023

Challenging Common Assumptions in Convex Reinforcement Learning

Mirco Mutti, Riccardo De Santi, Piersilvio De Bartolomeis +1

The classic Reinforcement Learning (RL) formulation concerns the maximization of a scalar reward function. More recently, convex RL has been introduced to extend the RL formulation…

cs.LG2022

Smoothing Policies and Safe Policy Gradients

Matteo Papini, Matteo Pirotta, Marcello Restelli

Policy Gradient (PG) algorithms are among the best candidates for the much-anticipated applications of reinforcement learning to real-world control tasks, such as robotics. However…

cs.LG2016

Unimodal Thompson Sampling for Graph-Structured Arms

Stefano Paladino, Francesco Trovò, Marcello Restelli +1

We study, to the best of our knowledge, the first Bayesian algorithm for unimodal Multi-Armed Bandit (MAB) problems with graph structure. In this setting, each arm corresponds to a…

cs.LG2025

Enhancing Diversity in Parallel Agents: A Maximum State Entropy Exploration Story

Vincenzo De Paola, Riccardo Zamboni, Mirco Mutti +1

Parallel data collection has redefined Reinforcement Learning (RL), unlocking unprecedented efficiency and powering breakthroughs in large-scale real-world applications. In this pa…

cs.LG2024

Interpetable Target-Feature Aggregation for Multi-Task Learning based on Bias-Variance Analysis

Paolo Bonetti, Alberto Maria Metelli, Marcello Restelli

Multi-task learning (MTL) is a powerful machine learning paradigm designed to leverage shared knowledge across tasks to improve generalization and performance. Previous works have…

cs.LG2020

Online Joint Bid/Daily Budget Optimization of Internet Advertising Campaigns

Alessandro Nuara, Francesco Trovò, Nicola Gatti +1

Pay-per-click advertising includes various formats (\emph{e.g.}, search, contextual, social) with a total investment of more than 200 billion USD per year worldwide. An advertiser…

cs.LG2026

Actor-Critic with Active Importance Sampling

Majid Molaei, Gabor Paczolay, Matteo Papini +2

This paper introduces the Active-Importance-Sampling Actor-Critic (AISAC) algorithm, an extension of the Actor-Critic framework for reducing variance in policy gradient estimation.…

cs.LG2023

Towards Theoretical Understanding of Inverse Reinforcement Learning

Alberto Maria Metelli, Filippo Lazzati, Marcello Restelli

Inverse reinforcement learning (IRL) denotes a powerful family of algorithms for recovering a reward function justifying the behavior demonstrated by an expert agent. A well-known…

cs.LG2022

Reward-Free Policy Space Compression for Reinforcement Learning

Mirco Mutti, Stefano Del Col, Marcello Restelli

In reinforcement learning, we encode the potential behaviors of an agent interacting with an environment into an infinite set of policies, the policy space, typically represented b…

cs.LG2018

Policy Optimization via Importance Sampling

Alberto Maria Metelli, Matteo Papini, Francesco Faccio +1

Policy optimization is an effective reinforcement learning approach to solve continuous control tasks. Recent achievements have shown that alternating online and offline optimizati…

cs.LG2025

"So, Tell Me About Your Policy...": Distillation of interpretable policies from Deep Reinforcement Learning agents

Giovanni Dispoto, Paolo Bonetti, Marcello Restelli

Recent advances in Reinforcement Learning (RL) largely benefit from the inclusion of Deep Neural Networks, boosting the number of novel approaches proposed in the field of Deep Rei…

cs.LG2024

Statistical Analysis of Policy Space Compression Problem

Majid Molaei, Marcello Restelli, Alberto Maria Metelli +1

Policy search methods are crucial in reinforcement learning, offering a framework to address continuous state-action and partially observable problems. However, the complexity of e…

cs.LG2024

Inverse Reinforcement Learning with Sub-optimal Experts

Riccardo Poiani, Gabriele Curti, Alberto Maria Metelli +1

Inverse Reinforcement Learning (IRL) techniques deal with the problem of deducing a reward function that explains the behavior of an expert agent who is assumed to act optimally in…

cs.LG2023

Nonlinear Feature Aggregation: Two Algorithms driven by Theory

Paolo Bonetti, Alberto Maria Metelli, Marcello Restelli

Many real-world machine learning applications are characterized by a huge number of features, leading to computational and memory issues, as well as the risk of overfitting. Ideall…

cs.LG2026

From Parameters to Behaviors: Unsupervised Compression of the Policy Space

Davide Tenedini, Riccardo Zamboni, Mirco Mutti +1

Despite its recent successes, Deep Reinforcement Learning (DRL) is notoriously sample-inefficient. We argue that this inefficiency stems from the standard practice of optimizing po…

cs.LG2026

How Log-Barrier Helps Exploration in Policy Optimization

Leonardo Cesani, Matteo Papini, Marcello Restelli

Recently, it has been shown that the Stochastic Gradient Bandit (SGB) algorithm converges to a globally optimal policy with a constant learning rate. However, these guarantees rely…

cs.LG2025

Optimizing Energy Management of Smart Grid using Reinforcement Learning aided by Surrogate models built using Physics-informed Neural Networks

Julen Cestero, Carmine Delle Femine, Kenji S. Muro +2

Optimizing the energy management within a smart grids scenario presents significant challenges, primarily due to the complexity of real-world systems and the intricate interactions…

cs.LG2023

Truncating Trajectories in Monte Carlo Reinforcement Learning

Riccardo Poiani, Alberto Maria Metelli, Marcello Restelli

In Reinforcement Learning (RL), an agent acts in an unknown environment to maximize the expected cumulative discounted sum of an external reward signal, i.e., the expected return.…

cs.LG2022

Multi-Armed Bandit Problem with Temporally-Partitioned Rewards: When Partial Feedback Counts

Giulia Romano, Andrea Agostini, Francesco Trovò +2

There is a rising interest in industrial online applications where data becomes available sequentially. Inspired by the recommendation of playlists to users where their preferences…

cs.LG2023

Wasserstein Actor-Critic: Directed Exploration via Optimism for Continuous-Actions Control

Amarildo Likmeta, Matteo Sacco, Alberto Maria Metelli +1

Uncertainty quantification has been extensively used as a means to achieve efficient directed exploration in Reinforcement Learning (RL). However, state-of-the-art methods for cont…

cs.AI2021

A Practical Guide to Multi-Objective Reinforcement Learning and Planning

Conor F. Hayes, Roxana Rădulescu, Eugenio Bargiacchi +15

Real-world decision-making tasks are generally complex, requiring trade-offs between multiple, often conflicting, objectives. Despite this, the majority of research in reinforcemen…

cs.LG2020

An Asymptotically Optimal Primal-Dual Incremental Algorithm for Contextual Linear Bandits

Andrea Tirinzoni, Matteo Pirotta, Marcello Restelli +1

In the contextual linear bandit setting, algorithms built on the optimism principle fail to exploit the structure of the problem and have been shown to be asymptotically suboptimal…

cs.LG2022

The Importance of Non-Markovianity in Maximum State Entropy Exploration

Mirco Mutti, Riccardo De Santi, Marcello Restelli

In the maximum state entropy exploration framework, an agent interacts with a reward-free environment to learn a policy that maximizes the entropy of the expected state visitations…