6 citations · 6 across the 1 of their papers we have counts for
1 paper
Cunchen Hu, Heyang Huang, Liangliang Xu +9
Transformer-based large language model (LLM) inference serving is now the backbone of many cloud services. LLM inference consists of a prefill phase and a decode phase. However, ex…