1 paper · 1 filter
Omar Basit, Yunzhao Liu, Z. Jonny Kong +1
Prefill/decode disaggregation is increasingly adopted in LLM serving to improve the latency-throughput tradeoff and meet strict TTFT and TPOT SLOs. However, LLM inference remains e…