Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
InstCache: A Predictive Cache for LLM Serving
Longwei Zou, Yan Liu, Jiamu Kang +3
The revolutionary capabilities of Large Language Models (LLMs) are attracting rapidly growing popularity and leading to soaring user requests to inference serving systems. Caching…
cs.CL2024
CQIL: Inference Latency Optimization with Concurrent Computation of Quasi-Independent Layers
Longwei Zou, Qingyang Wang, Han Zhao +3
The fast-growing large scale language models are delivering unprecedented performance on almost all natural language processing tasks. However, the effectiveness of large language…