47 citations · 49 across the 2 of their papers we have counts for
2 papers
cs.LG2023★ 47 cited
FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU
Ying Sheng, Lianmin Zheng, Binhang Yuan +11
The high computational and memory requirements of large language model (LLM) inference make it feasible only with multiple high-end accelerators. Motivated by the emerging demand f…
cs.AR2021★ 2 cited
Dual-side Sparse Tensor Core
Yang Wang, Chen Zhang, Zhiqiang Xie +3
Leveraging sparsity in deep neural network (DNN) models is promising for accelerating model inference. Yet existing GPUs can only leverage the sparsity from weights but not activat…