activity
20242026
collaborators
Showing cs.LGShow all

34 papers · 1 filter

cs.LG2026

Unsupervised Behavioral Compression: Learning Low-Dimensional Policy Manifolds through State-Occupancy Matching

Andrea Fraschini, Davide Tenedini, Riccardo Zamboni +2

Deep Reinforcement Learning (DRL) is widely recognized as sample-inefficient, a limitation attributable in part to the high dimensionality and substantial functional redundancy inh…

cs.LG2026

Online Market Making and the Value of Observing the Order Book

Davide Maran, Marcello Restelli

We study an online market-making problem in which a learner sequentially posts bid and ask prices for a single asset while interacting with traders holding private valuations. Unli…

cs.LG2026

Actor-Critic with Active Importance Sampling

Majid Molaei, Gabor Paczolay, Matteo Papini +2

This paper introduces the Active-Importance-Sampling Actor-Critic (AISAC) algorithm, an extension of the Actor-Critic framework for reducing variance in policy gradient estimation.…

cs.LG2026

How Log-Barrier Helps Exploration in Policy Optimization

Leonardo Cesani, Matteo Papini, Marcello Restelli

Recently, it has been shown that the Stochastic Gradient Bandit (SGB) algorithm converges to a globally optimal policy with a constant learning rate. However, these guarantees rely…

cs.LG2026

Finite Sample Bounds for Non-Parametric Regression: Optimal Sample Efficiency and Space Complexity

Davide Maran, Marcello Restelli

We address the problem of learning an unknown smooth function and its derivatives from noisy pointwise evaluations under the supremum norm. While classical nonparametric regression…

cs.LG2026

From Parameters to Behaviors: Unsupervised Compression of the Policy Space

Davide Tenedini, Riccardo Zamboni, Mirco Mutti +1

Despite its recent successes, Deep Reinforcement Learning (DRL) is notoriously sample-inefficient. We argue that this inefficiency stems from the standard practice of optimizing po…