1 paper
Amna Masood, Pratishtha Gaur, Nuwan Jayasena
Two widely adopted techniques for LLM inference serving systems today are hybrid batching and disaggregated serving. A hybrid batch combines prefill and decode tokens of different…