47 citations · 67 across the 3 of their papers we have counts for
3 papers
cs.LG2023★ 19 cited
Deja Vu: Contextual Sparsity for Efficient LLMs at Inference Time
Zichang Liu, Jue Wang, Tri Dao +8
Large language models (LLMs) with hundreds of billions of parameters have sparked a new wave of exciting AI applications. However, they are computationally expensive at inference t…
cs.LG2023★ 1 cited
Auto-Differentiation of Relational Computations for Very Large Scale Machine Learning
Yuxin Tang, Zhimin Ding, Dimitrije Jankov +3
The relational data model was designed to facilitate large-scale data management and analytics. We consider the problem of how to differentiate computations expressed relationally.…
cs.LG2023★ 47 cited
FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU
Ying Sheng, Lianmin Zheng, Binhang Yuan +11
The high computational and memory requirements of large language model (LLM) inference make it feasible only with multiple high-end accelerators. Motivated by the emerging demand f…