1 paper · 1 filter
Zongze Li, Jingyu Liu, Zhen Xu +3
Prefill-Decode (PD) disaggregation has become the standard architecture for modern LLM inference engines, which alleviates the interference of two distinctive workloads. With the g…