most citedDeja Vu: Contextual Sparsity for Efficient LLMs at Inference Time

19 citations · 31 across the 6 of their papers we have counts for

collaborators

8 papers

cs.LG202319 cited

Deja Vu: Contextual Sparsity for Efficient LLMs at Inference Time

Zichang Liu, Jue Wang, Tri Dao +8

Large language models (LLMs) with hundreds of billions of parameters have sparked a new wave of exciting AI applications. However, they are computationally expensive at inference t…

cs.LG20231 cited

How to Protect Copyright Data in Optimization of Large Language Models?

Timothy Chu, Zhao Song, Chiwun Yang

Large language models (LLMs) and generative AI have played a transformative role in computer research and applications. Controversy has arisen as to whether these models output cop…

cs.LG2023

Clustered Linear Contextual Bandits with Knapsacks

Yichuan Deng, Michalis Mamakos, Zhao Song

In this work, we study clustered contextual bandits where rewards and resource consumption are the outcomes of cluster-specific linear models. The arms are divided in clusters, wit…

cs.LG20236 cited

GradientCoin: A Peer-to-Peer Decentralized Large Language Models

Yeqi Gao, Zhao Song, Junze Yin

Since 2008, after the proposal of a Bitcoin electronic cash system, Bitcoin has fundamentally changed the economic system over the last decade. Since 2022, large language models (L…

cs.LG20232 cited

Convergence of Two-Layer Regression with Nonlinear Units

Yichuan Deng, Zhao Song, Shenghao Xie

Large language models (LLMs), such as ChatGPT and GPT4, have shown outstanding performance in many human life task. Attention computation plays an important role in training LLMs.…

cs.LG20233 cited

Zero-th Order Algorithm for Softmax Attention Optimization

Yichuan Deng, Zhihang Li, Sridhar Mahadevan +1

Large language models (LLMs) have brought about significant transformations in human society. Among the crucial computations in LLMs, the softmax unit holds great importance. Its h…