3 papers
cs.AR2026
Adaptive KV Cache Reuse for Fast Long-Context LLM Serving
Fei li, Song Liu, Yan Liu +4
In long-context Large Language Model (LLM) inference, the Time-To-First-Token (TTFT) latency incurred by the prefill stage has become the foremost bottleneck limiting interactive p…
cs.LG2025
KVmix: Gradient-Based Layer Importance-Aware Mixed-Precision Quantization for KV Cache
Fei Li, Song Liu, Weiguo Wu +2
The high memory demands of the Key-Value (KV) Cache during the inference of Large Language Models (LLMs) severely restrict their deployment in resource-constrained platforms. Quant…
cs.DC2014
Scalable Hierarchical Scheduling for Malleable Parallel Jobs on Multiprocessor-based Systems
Yangjie Cao, Hongyang Sun, Depei Qian +1
The proliferation of multi-core and multiprocessor-based computer systems has led to explosive development of parallel applications and hence the need for efficient schedulers. In…