2 papers
cs.AR2026
Resource-Efficient Speculative Decoding for Long-Context LLM Serving
Fei Li, Song Liu, Shiqiang Nie +2
Speculative decoding reduces sequential Target model calls by verifying multiple tokens from the Draft model in parallel. Yet KV Cache growth limits long-context serving under cons…
cs.LG2025
KVmix: Gradient-Based Layer Importance-Aware Mixed-Precision Quantization for KV Cache
Fei Li, Song Liu, Weiguo Wu +2
The high memory demands of the Key-Value (KV) Cache during the inference of Large Language Models (LLMs) severely restrict their deployment in resource-constrained platforms. Quant…