2 citations · 2 across the 1 of their papers we have counts for
1 paper
Jiaao He, Jidong Zhai
Cost of serving large language models (LLM) is high, but the expensive and scarce GPUs are poorly efficient when generating tokens sequentially, unless the batch of sequences is en…