distributed cache 1hybrid sliding window attention 1kvcache optimization 1mixture-of-experts 1multimodal inference 1
From the 1 of 8 linked papers with an AI index.
Showing cs.DCShow all
2 papers · 1 filter
cs.DC2026
SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving
Pol G. Recasens, Ferran Agullo, Yue Zhu +3
Large language model (LLM) serving commonly increases batch size to improve throughput, but performance eventually reaches a deployment-dependent plateau beyond which larger batche…
cs.DC2025
Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference
Pol G. Recasens, Ferran Agullo, Yue Zhu +5
Large language models have been widely adopted across different tasks, but their auto-regressive generation nature often leads to inefficient resource utilization during inference.…