5 papers
Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning
Joseph Lazzaro, Alessio Russo, Aldo Pacchiano
In this work we study the Best Policy Identification (BPI) problem in online, tabular Reinforcement Learning. This is an active sequential hypothesis testing problem in which the l…
A Finite Time Analysis of Thompson Sampling for Bayesian Optimization with Preferential Feedback
Joseph Lazzaro, Davide Buffelli, Da-shan Shiu +1
Preference feedback, in the form of pairwise comparisons rather than scalar scores, has seen increasing use in applications such as human-, laboratory-, and expert-in-the-loop desi…
Locally Differentially Private Thresholding Bandits
Annalisa Barbara, Joseph Lazzaro, Ciara Pike-Burke
This work investigates the impact of ensuring local differential privacy in the thresholding bandit problem. We consider both the fixed budget and fixed confidence settings. We pro…
Fixed-Confidence Multiple Change Point Identification under Bandit Feedback
Joseph Lazzaro, Ciara Pike-Burke
Piecewise constant functions describe a variety of real-world phenomena in domains ranging from chemistry to manufacturing. In practice, it is often required to confidently identif…
Fixed-Budget Change Point Identification in Piecewise Constant Bandits
Joseph Lazzaro, Ciara Pike-Burke
We study the piecewise constant bandit problem where the expected reward is a piecewise constant function with one change point (discontinuity) across the action space and…