2 papers
cs.LG2026
On the Benefits of Free Exploration for Regret Minimization in Multi-Armed Bandits
Yunlong Hou, Zixin Zhong, Vincent Y. F. Tan
We study a stochastic multi-armed bandit problem where an agent is granted a free exploration budget before regret accumulates, a setting not captured by the classic regret minimiz…
cs.LG2024
Almost Minimax Optimal Best Arm Identification in Piecewise Stationary Linear Bandits
Yunlong Hou, Vincent Y. F. Tan, Zixin Zhong
We propose a {\em novel} piecewise stationary linear bandit (PSLB) model, where the environment randomly samples a context from an unknown probability distribution at each changepo…