1 paper
Shuowei Jin, Xueshen Liu, Jiaxin Shan +4
As LLM inference shifts to multi-tenant GPU clusters, co-batching improves throughput but obscures per-tenant usage and limits control. Enabling fractional sharing of the inference…