1 paper
Yipin Guo, Siddharth Joshi
Contemporary systems serving large language models (LLMs) have adopted prefill-decode disaggregation to load-balance between the compute-bound prefill phase and the memory-bound de…