3 papers
cs.LG2026
Breaking the Total Variance Barrier: Sharp Sample Complexity for Linear Heteroscedastic Bandits with Fixed Action Set
Heyang Zhao, Tianyuan Jin, Weixin Wang +3
Recent years have witnessed increasing interests in tackling heteroscedastic noise in bandits and reinforcement learning. In these works, the cumulative variance of the noise $Î=…
cs.LG2026
Best Arm Identification with Minimal Regret
Junwen Yang, Vincent Y. F. Tan, Tianyuan Jin
Motivated by real-world applications that necessitate responsible experimentation, we introduce the problem of best arm identification (BAI) with minimal regret. This variant of th…
cs.LG2026
Asymptotically and Minimax Optimal Regret Bounds for Multi-Armed Bandits with Abstention
Junwen Yang, Tianyuan Jin, Vincent Y. F. Tan
We introduce a novel extension of the canonical multi-armed bandit problem that incorporates an additional strategic innovation: abstention. In this enhanced framework, the agent i…