1 paper
D. Balamurugan, Thomas W. Bush
Serving large language models (LLMs) on a shared, heterogeneous GPU cluster requires users and operators to select the GPU type, tensor-parallel degree, and precision before commit…