Showing stat.MLShow all
2 papers · 1 filter
stat.ML2026
Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning
Joseph Lazzaro, Alessio Russo, Aldo Pacchiano
In this work we study the Best Policy Identification (BPI) problem in online, tabular Reinforcement Learning. This is an active sequential hypothesis testing problem in which the l…
stat.ML2025
Pure Exploration with Feedback Graphs
Alessio Russo, Yichen Song, Aldo Pacchiano
We study the sample complexity of pure exploration in an online learning problem with a feedback graph. This graph dictates the feedback available to the learner, covering scenario…