6 papers
Pure Exploration for a Good Policy in Reinforcement Learning with Bandit Feedback
Zitian Li, Wang Chi Cheung
Pure exploration in episodic Reinforcement Learning has primarily focused on Best Policy Identification (BPI), which seeks to identify a (near)-optimal policy with high confidence.…
Closing the Gap on the Sample Complexity of 1-Identification
Zitian Li, Wang Chi Cheung
The 1-identification problem is a fundamental pure-exploration problem in multi-armed bandits. An agent aims to determine whether there exists an arm whose mean reward exceeds a kn…
Learning with a Budget: Identifying the Best Arm with Resource Constraints
Zitian Li, Wang Chi Cheung
In many applications, evaluating the effectiveness of different alternatives comes with varying costs or resource usage. Motivated by such heterogeneity, we study the Best Arm Iden…
Episodic Contextual Bandits with Knapsacks under Conversion Models
Wang Chi Cheung, Zitian Li
We study an online setting, where a decision maker (DM) interacts with contextual bandit-with-knapsack (BwK) instances in repeated episodes. These episodes start with different res…
Near Optimal Non-asymptotic Sample Complexity of 1-Identification
Zitian Li, Wang Chi Cheung
Motivated by an open direction in existing literature, we study the 1-identification problem, a fundamental multi-armed bandit formulation on pure exploration. The goal is to deter…
Best Arm Identification with Resource Constraints
Zitian Li, Wang Chi Cheung
Motivated by the cost heterogeneity in experimentation across different alternatives, we study the Best Arm Identification with Resource Constraints (BAIwRC) problem. The agent aim…