2 papers
cs.LG2026
Not-a-Bandit: Provably No-Regret Drafter Selection in Speculative Decoding for LLMs
Hongyi Liu, Jiaji Huang, Zhen Jia +2
Speculative decoding is widely used in accelerating large language model (LLM) inference. In this work, we focus on the online draft model selection problem in speculative decoding…
cs.LG2025
ProxSparse: Regularized Learning of Semi-Structured Sparsity Masks for Pretrained LLMs
Hongyi Liu, Rajarshi Saha, Zhen Jia +5
Large Language Models (LLMs) have demonstrated exceptional performance in natural language processing tasks, yet their massive size makes serving them inefficient and costly. Semi-…