1 paper
Jiaxi Li, Yue Zhu, Eun Kyung Lee +1
Different from traditional Large Language Model (LLM) serving that colocates the prefill and decode stages on the same GPU, disaggregated serving dedicates distinct GPUs to prefill…