3 papers
cs.LG2026
Breaking the Ice: Analyzing Cold Start Latency in vLLM
Huzaifa Shaaban Kabakibo, Animesh Trivedi, Lin Wang
As scalable inference services become popular, the cold start latency of an inference engine becomes important. Today, vLLM has evolved into the de-facto inference engine of choice…
cs.DC2026
Kavier: Exploring Performance, Sustainability, and Efficiency of LLM Ecosystems under Inference through Cache-Aware Discrete-Event Simulation
Radu Nicolae, Alexandru Iosup, Animesh Trivedi +1
Large Language Models (LLMs) are widely used by our increasingly digitalized society, but raise sustainability, performance, and financial concerns, especially as inference workloa…
cs.DC2024
Exploring GPU-to-GPU Communication: Insights into Supercomputer Interconnects
Daniele De Sensi, Lorenzo Pichetti, Flavio Vella +11
Multi-GPU nodes are increasingly common in the rapidly evolving landscape of exascale supercomputers. On these systems, GPUs on the same node are connected through dedicated networ…