8 citations · 8 across the 4 of their papers we have counts for
1 paper · 1 filter
Danying Ge, Jianhua Gao, Qizhi Jiang +2
Speculative decoding, which combines a draft model with a target model, has emerged as an effective approach to accelerate large language model (LLM) inference. However, existing m…