2 papers
cs.DC2025
ESS: An Offload-Centric Latent-Cache Management Architecture for DeepSeek-V3.2-Exp
Xinhang Chen, Chao Zhang, Jiahuan He +9
DeepSeek-V3.2-Exp introduces a sparse attention mechanism that significantly reduces inference latency in long-context scenarios. Although the overall throughput has improved great…
cs.DB2024
The CAP Principle for LLM Serving: A Survey of Long-Context Large Language Model Serving
Pai Zeng, Zhenyu Ning, Jieru Zhao +5
We survey the large language model (LLM) serving area to understand the intricate dynamics between cost-efficiency and accuracy, which is magnified by the growing need for longer c…