activity
20132026
most citedSharing Knowledge in Multi-Task Deep Reinforcement Learning

62 citations · 151 across the 82 of their papers we have counts for

collaborators
Showing 2024 · cs.LGShow all

15 papers · 2 filters

cs.LG2024

Finite Sample Bounds for Non-Parametric Regression: Optimal Sample Efficiency and Space Complexity

Davide Maran, Marcello Restelli

We address the problem of learning an unknown smooth function and its derivatives from noisy pointwise evaluations under the supremum norm. While classical nonparametric regression…

cs.LG2024

Statistical Analysis of Policy Space Compression Problem

Majid Molaei, Marcello Restelli, Alberto Maria Metelli +1

Policy search methods are crucial in reinforcement learning, offering a framework to address continuous state-action and partially observable problems. However, the complexity of e…

cs.LG2024

Local Linearity: the Key for No-regret Reinforcement Learning in Continuous MDPs

Davide Maran, Alberto Maria Metelli, Matteo Papini +1

Achieving the no-regret property for Reinforcement Learning (RL) problems in continuous state and action-space environments is one of the major open problems in the field. Existing…

cs.LG2024★ 1 cited

Truncating Trajectories in Monte Carlo Policy Evaluation: an Adaptive Approach

Riccardo Poiani, Nicole Nobili, Alberto Maria Metelli +1

Policy evaluation via Monte Carlo (MC) simulation is at the core of many MC Reinforcement Learning (RL) algorithms (e.g., policy gradient methods). In this context, the designer of…

cs.LG2024

Efficient Learning of POMDPs with Known Observation Model in Average-Reward Setting

Alessio Russo, Alberto Maria Metelli, Marcello Restelli

Dealing with Partially Observable Markov Decision Processes is notably a challenging task. We face an average-reward infinite-horizon POMDP setting with an unknown transition model…

cs.LG2024★ 1 cited

The Limits of Pure Exploration in POMDPs: When the Observation Entropy is Enough

Riccardo Zamboni, Duilio Cirino, Marcello Restelli +1

The problem of pure exploration in Markov decision processes has been cast as maximizing the entropy over the state distribution induced by the agent's policy, an objective that ha…