1 citations · 1 across the 1 of their papers we have counts for
2 papers
cs.LG2024
Is the GPU Half-Empty or Half-Full? Practical Scheduling Techniques for LLMs
Ferdi Kossmann, Bruce Fontaine, Daya Khudia +2
Serving systems for Large Language Models (LLMs) improve throughput by processing several requests concurrently. However, multiplexing hardware resources between concurrent request…
cs.DC2024★ 1 cited
CascadeServe: Unlocking Model Cascades for Inference Serving
Ferdi Kossmann, Ziniu Wu, Alex Turk +3
Machine learning (ML) models are increasingly deployed to production, calling for efficient inference serving systems. Efficient inference serving is complicated by two challenges:…