27 papers
Unsupervised Behavioral Compression: Learning Low-Dimensional Policy Manifolds through State-Occupancy Matching
Andrea Fraschini, Davide Tenedini, Riccardo Zamboni +2
Deep Reinforcement Learning (DRL) is widely recognized as sample-inefficient, a limitation attributable in part to the high dimensionality and substantial functional redundancy inh…
Bridging Rested and Restless Bandits with Graph-Triggering: Rising and Rotting
Gianmarco Genalti, Marco Mussi, Nicola Gatti +3
Rested and Restless Bandits are two well-known bandit settings that are useful to model real-world sequential decision-making problems in which the expected reward of an arm evolve…
Online Market Making and the Value of Observing the Order Book
Davide Maran, Marcello Restelli
We study an online market-making problem in which a learner sequentially posts bid and ask prices for a single asset while interacting with traders holding private valuations. Unli…
Actor-Critic with Active Importance Sampling
Majid Molaei, Gabor Paczolay, Matteo Papini +2
This paper introduces the Active-Importance-Sampling Actor-Critic (AISAC) algorithm, an extension of the Actor-Critic framework for reducing variance in policy gradient estimation.…
How Log-Barrier Helps Exploration in Policy Optimization
Leonardo Cesani, Matteo Papini, Marcello Restelli
Recently, it has been shown that the Stochastic Gradient Bandit (SGB) algorithm converges to a globally optimal policy with a constant learning rate. However, these guarantees rely…
Finite Sample Bounds for Non-Parametric Regression: Optimal Sample Efficiency and Space Complexity
Davide Maran, Marcello Restelli
We address the problem of learning an unknown smooth function and its derivatives from noisy pointwise evaluations under the supremum norm. While classical nonparametric regression…