25 citations · 45 across the 9 of their papers we have counts for
Showing 2026 · cs.LGShow all
2 papers · 2 filters
cs.LG2026
Pure Exploration for a Good Policy in Reinforcement Learning with Bandit Feedback
Zitian Li, Wang Chi Cheung
Pure exploration in episodic Reinforcement Learning has primarily focused on Best Policy Identification (BPI), which seeks to identify a (near)-optimal policy with high confidence.…
cs.LG2026
Closing the Gap on the Sample Complexity of 1-Identification
Zitian Li, Wang Chi Cheung
The 1-identification problem is a fundamental pure-exploration problem in multi-armed bandits. An agent aims to determine whether there exists an arm whose mean reward exceeds a kn…