activity
20162024
most citedFlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU

47 citations · 93 across the 8 of their papers we have counts for

collaborators
Showing cs.LGShow all

5 papers · 1 filter

cs.LG20234 cited

Laughing Hyena Distillery: Extracting Compact Recurrences From Convolutions

Stefano Massaroli, Michael Poli, Daniel Y. Fu +11

Recent advances in attention-free sequence models rely on convolutions as alternatives to the attention operator at the core of Transformers. In particular, long convolution sequen…

cs.LG202319 cited

Deja Vu: Contextual Sparsity for Efficient LLMs at Inference Time

Zichang Liu, Jue Wang, Tri Dao +8

Large language models (LLMs) with hundreds of billions of parameters have sparked a new wave of exciting AI applications. However, they are computationally expensive at inference t…

cs.LG202347 cited

FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU

Ying Sheng, Lianmin Zheng, Binhang Yuan +11

The high computational and memory requirements of large language model (LLM) inference make it feasible only with multiple high-end accelerators. Motivated by the emerging demand f…

cs.LG2023

Sample-efficient Surrogate Model for Frequency Response of Linear PDEs using Self-Attentive Complex Polynomials

Andrew Cohen, Weiping Dou, Jiang Zhu +7

Linear Partial Differential Equations (PDEs) govern the spatial-temporal dynamics of physical systems that are essential to building modern technology. When working with linear PDE…

cs.LG20217 cited

Pixelated Butterfly: Simple and Efficient Sparse training for Neural Network Models

Tri Dao, Beidi Chen, Kaizhao Liang +4

Overparameterized neural networks generalize well but are expensive to train. Ideally, one would like to reduce their computational cost while retaining their generalization benefi…