9 papers
Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning
Joseph Lazzaro, Alessio Russo, Aldo Pacchiano
In this work we study the Best Policy Identification (BPI) problem in online, tabular Reinforcement Learning. This is an active sequential hypothesis testing problem in which the l…
Receding-Horizon Control via Drifting Models
Daniele Foffano, Alessio Russo, Alexandre Proutiere
We study the problem of trajectory optimization in settings where the system dynamics are unknown and it is not possible to simulate trajectories through a surrogate model. When an…
In-Context Learning for Pure Exploration
Alessio Russo, Ryan Welch, Aldo Pacchiano
We study the problem active sequential hypothesis testing, also known as pure exploration: given a new task, the learner adaptively collects data from the environment to efficientl…
In-Context Pure Exploration in Continuous Decision Spaces
Alessio Russo, Yin-Ching Lee, Ryan Welch +1
In active sequential testing, also termed pure exploration, a learner is tasked with the goal to adaptively acquire information so as to identify an unknown ground-truth hypothesis…
Achieving Regret in Average-Reward POMDPs with Known Observation Models
Alessio Russo, Alberto Maria Metelli, Marcello Restelli
We tackle average-reward infinite-horizon POMDPs with an unknown transition model but a known observation model, a setting that has been previously addressed in two limiting ways:…
Adaptive Exploration for Multi-Reward Multi-Policy Evaluation
Alessio Russo, Aldo Pacchiano
We study the policy evaluation problem in an online multi-reward multi-policy discounted setting, where multiple reward functions must be evaluated simultaneously for different pol…