Showing cs.ARShow all
2 papers · 1 filter
cs.AR2026
Resource-Efficient Speculative Decoding for Long-Context LLM Serving
Fei Li, Song Liu, Shiqiang Nie +2
Speculative decoding reduces sequential Target model calls by verifying multiple tokens from the Draft model in parallel. Yet KV Cache growth limits long-context serving under cons…
cs.AR2026
Adaptive KV Cache Reuse for Fast Long-Context LLM Serving
Fei li, Song Liu, Yan Liu +4
In long-context Large Language Model (LLM) inference, the Time-To-First-Token (TTFT) latency incurred by the prefill stage has become the foremost bottleneck limiting interactive p…