1 paper
Yiding Feng, Zonghan Yang, Yuhao Zhang
Large Language Model (LLM) inference presents a unique scheduling challenge due to the Key-Value (KV) cache, where a job's memory footprint grows linearly with the number of decode…