1 paper
Jaehun Lee, In-Jun Jung, Joo-Young Kim
Large language model (LLM) serving increasingly combines prefill-decode (PD) disaggregation with tensor parallelism (TP) to support large models and long contexts. In conventional…