15 citations · 17 across the 2 of their papers we have counts for
2 papers
cs.AI2024★ 2 cited
Fast On-device LLM Inference with NPUs
Daliang Xu, Hao Zhang, Liming Yang +4
On-device inference for Large Language Models (LLMs), driven by increasing privacy concerns and advancements of mobile-sized models, has gained significant interest. However, even…
cs.DC2024★ 15 cited
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
Yinmin Zhong, Shengyu Liu, Junda Chen +5
DistServe improves the performance of large language models (LLMs) serving by disaggregating the prefill and decoding computation. Existing LLM serving systems colocate the two pha…