1 paper
Jianxing Qin, Alexander Du, Danfeng Zhang +2
LLM inference systems may vary batch composition, prompt chunking, prefill/decode execution, and KV-cache reuse, eviction, or recomputation. These optimizations should not affect s…