2 papers
cs.CL2026
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression
DeepSeek-AI, :, Anyi Xu +585
The widespread adoption of long-horizon agents has made model workloads increasingly input-heavy. Although prior work has substantially reduced the cost of long-context computation…
cs.DC2026
Measurement-Driven Diagnosis and Mitigation of Host-CPU Co-location Interference in Single-GPU LLM Serving on a Multi-GPU Server
Guanjie Cheng, Guowei Li, Yingying Wen +3
Host CPUs in GPU servers are often under-used during LLM inference. Co-locating CPU workloads can improve resource use, but it can also seriously hurt serving quality. Existing wor…