1 paper
Guanjie Cheng, Guowei Li, Yingying Wen +3
Host CPUs in GPU servers are often under-used during LLM inference. Co-locating CPU workloads can improve resource use, but it can also seriously hurt serving quality. Existing wor…