activity
20182026
most citedThe Limits of Pure Exploration in POMDPs: When the Observation Entropy is Enough

1 citations · 1 across the 13 of their papers we have counts for

collaborators
Showing cs.LGShow all

20 papers · 1 filter

cs.LG2026

Unsupervised Behavioral Compression: Learning Low-Dimensional Policy Manifolds through State-Occupancy Matching

Andrea Fraschini, Davide Tenedini, Riccardo Zamboni +2

Deep Reinforcement Learning (DRL) is widely recognized as sample-inefficient, a limitation attributable in part to the high dimensionality and substantial functional redundancy inh…

cs.LG2026

K-Myriad: Jump-starting reinforcement learning with unsupervised parallel agents

Vincenzo De Paola, Mirco Mutti, Riccardo Zamboni +1

Parallelization in Reinforcement Learning is typically employed to speed up the training of a single policy, where multiple workers collect experience from an identical sampling di…

cs.LG2025

From Parameters to Behaviors: Unsupervised Compression of the Policy Space

Davide Tenedini, Riccardo Zamboni, Mirco Mutti +1

Despite its recent successes, Deep Reinforcement Learning (DRL) is notoriously sample-inefficient. We argue that this inefficiency stems from the standard practice of optimizing po…

cs.LG2025

Enhancing Diversity in Parallel Agents: A Maximum State Entropy Exploration Story

Vincenzo De Paola, Riccardo Zamboni, Mirco Mutti +1

Parallel data collection has redefined Reinforcement Learning (RL), unlocking unprecedented efficiency and powering breakthroughs in large-scale real-world applications. In this pa…

cs.LG2025

State Entropy Regularization for Robust Reinforcement Learning

Yonatan Ashlag, Uri Koren, Mirco Mutti +3

State entropy regularization has empirically shown better exploration and sample complexity in reinforcement learning (RL). However, its theoretical guarantees have not been studie…

cs.LG2025

A Classification View on Meta Learning Bandits

Mirco Mutti, Jeongyeol Kwon, Shie Mannor +1

Contextual multi-armed bandits are a popular choice to model sequential decision-making. E.g., in a healthcare application we may perform various tests to asses a patient condition…