activity
20202023
most citedDeja Vu: Contextual Sparsity for Efficient LLMs at Inference Time

19 citations · 70 across the 9 of their papers we have counts for

collaborators

10 papers

cs.LG202319 cited

Deja Vu: Contextual Sparsity for Efficient LLMs at Inference Time

Zichang Liu, Jue Wang, Tri Dao +8

Large language models (LLMs) with hundreds of billions of parameters have sparked a new wave of exciting AI applications. However, they are computationally expensive at inference t…

cs.CL2023

RLCD: Reinforcement Learning from Contrastive Distillation for Language Model Alignment

Kevin Yang, Dan Klein, Asli Celikyilmaz +2

We propose Reinforcement Learning from Contrastive Distillation (RLCD), a method for aligning language models to follow principles expressed in natural language (e.g., to be more h…

cs.LG2023

Landscape Surrogate: Learning Decision Losses for Mathematical Optimization Under Partial Information

Arman Zharmagambetov, Brandon Amos, Aaron Ferber +3

Recent works in learning-integrated optimization have shown promise in settings where the optimization problem is only partially observed or where general-purpose optimizers perfor…

cs.LG20221 cited

EurNet: Efficient Multi-Range Relational Modeling of Spatial Multi-Relational Data

Minghao Xu, Yuanfan Guo, Yi Xu +3

Modeling spatial relationship in the data remains critical across many different tasks, such as image classification, semantic segmentation and protein structure understanding. Pre…

cs.CL202211 cited

Re3: Generating Longer Stories With Recursive Reprompting and Revision

Kevin Yang, Yuandong Tian, Nanyun Peng +1

We consider the problem of automatically generating longer stories of over two thousand words. Compared to prior work on shorter stories, long-range plot coherence and relevance ar…

cs.LG20227 cited

DreamShard: Generalizable Embedding Table Placement for Recommender Systems

Daochen Zha, Louis Feng, Qiaoyu Tan +6

We study embedding table placement for distributed recommender systems, which aims to partition and place the tables on multiple hardware devices (e.g., GPUs) to balance the comput…