collaborators

6 papers

cs.LG2026

Pure Exploration for a Good Policy in Reinforcement Learning with Bandit Feedback

Zitian Li, Wang Chi Cheung

Pure exploration in episodic Reinforcement Learning has primarily focused on Best Policy Identification (BPI), which seeks to identify a (near)-optimal policy with high confidence.…

cs.LG2026

Closing the Gap on the Sample Complexity of 1-Identification

Zitian Li, Wang Chi Cheung

The 1-identification problem is a fundamental pure-exploration problem in multi-armed bandits. An agent aims to determine whether there exists an arm whose mean reward exceeds a kn…

cs.LG2026

Learning with a Budget: Identifying the Best Arm with Resource Constraints

Zitian Li, Wang Chi Cheung

In many applications, evaluating the effectiveness of different alternatives comes with varying costs or resource usage. Motivated by such heterogeneity, we study the Best Arm Iden…

cs.LG2026

Episodic Contextual Bandits with Knapsacks under Conversion Models

Wang Chi Cheung, Zitian Li

We study an online setting, where a decision maker (DM) interacts with contextual bandit-with-knapsack (BwK) instances in repeated episodes. These episodes start with different res…

cs.LG2025

Near Optimal Non-asymptotic Sample Complexity of 1-Identification

Zitian Li, Wang Chi Cheung

Motivated by an open direction in existing literature, we study the 1-identification problem, a fundamental multi-armed bandit formulation on pure exploration. The goal is to deter…

cs.LG2025

Best Arm Identification with Resource Constraints

Zitian Li, Wang Chi Cheung

Motivated by the cost heterogeneity in experimentation across different alternatives, we study the Best Arm Identification with Resource Constraints (BAIwRC) problem. The agent aim…