13 papers
Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning
Joseph Lazzaro, Alessio Russo, Aldo Pacchiano
In this work we study the Best Policy Identification (BPI) problem in online, tabular Reinforcement Learning. This is an active sequential hypothesis testing problem in which the l…
Meet Me at the Arm: The Cooperative Multi-Armed Bandits Problem with Shareable Arms
Xinyi Hu, Aldo Pacchiano
We study the decentralized multi-player multi-armed bandits (MMAB) problem under a no-sensing setting, where each player receives only their own reward and obtains no information a…
In-Context Learning for Pure Exploration
Alessio Russo, Ryan Welch, Aldo Pacchiano
We study the problem active sequential hypothesis testing, also known as pure exploration: given a new task, the learner adaptively collects data from the environment to efficientl…
In-Context Pure Exploration in Continuous Decision Spaces
Alessio Russo, Yin-Ching Lee, Ryan Welch +1
In active sequential testing, also termed pure exploration, a learner is tasked with the goal to adaptively acquire information so as to identify an unknown ground-truth hypothesis…
Bayesian Online Model Selection
Aida Afshar, Yuke Zhang, Aldo Pacchiano
Online model selection in Bayesian bandits raises a fundamental exploration challenge: When an environment instance is sampled from a prior distribution, how can we design an adapt…
Principled Fine-tuning of LLMs from User-Edits: A Medley of Preference, Supervision, and Reward
Dipendra Misra, Aldo Pacchiano, Ta-Chung Chi +1
We study how to fine-tune LLMs using user-edit deployment data consisting of a set of context, an agent's response, and user edits. This deployment data is naturally generated by u…