6 papers
Bayesian Best-Arm Identification with Abstention: A Polynomial-to-Exponential Phase Transition
Yuqi Huang, Yunlong Hou, Vincent Y. F. Tan
We study the Bayesian fixed-budget best-arm identification problem in which a learner can abstain from making a terminal recommendation. Subject to an abstention budget , we an…
On the Benefits of Free Exploration for Regret Minimization in Multi-Armed Bandits
Yunlong Hou, Zixin Zhong, Vincent Y. F. Tan
We study a stochastic multi-armed bandit problem where an agent is granted a free exploration budget before regret accumulates, a setting not captured by the classic regret minimiz…
Demystifying the Slash Pattern in Attention: The Role of RoPE
Yuan Cheng, Fengzhuo Zhang, Yunlong Hou +5
Large Language Models (LLMs) often exhibit slash attention patterns, where attention scores concentrate along the -th sub-diagonal for some offset . These patterns play a k…
BanditSpec: Adaptive Speculative Decoding via Bandit Algorithms
Yunlong Hou, Fengzhuo Zhang, Cunxiao Du +6
Speculative decoding has emerged as a popular method to accelerate the inference of Large Language Models (LLMs) while retaining their superior text generation performance. Previou…
Enhancing Long Video Generation Consistency without Tuning
Xingyao Li, Fengzhuo Zhang, Jiachun Pan +3
Despite the considerable progress achieved in the long video generation problem, there is still significant room to improve the consistency of the generated videos, particularly in…
Almost Minimax Optimal Best Arm Identification in Piecewise Stationary Linear Bandits
Yunlong Hou, Vincent Y. F. Tan, Zixin Zhong
We propose a {\em novel} piecewise stationary linear bandit (PSLB) model, where the environment randomly samples a context from an unknown probability distribution at each changepo…