activity
20182022
most citedLoRA: Low-Rank Adaptation of Large Language Models

2.5k citations · 2.6k across the 4 of their papers we have counts for

collaborators

7 papers

cs.LG202222 cited

Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Greg Yang, Edward J. Hu, Igor Babuschkin +7

Hyperparameter (HP) tuning in deep learning is an expensive process, prohibitively so for neural networks (NNs) with billions of parameters. We show that, in the recently discovere…

cs.CL202152 cited

Guided Generation of Cause and Effect

Zhongyang Li, Xiao Ding, Ting Liu +2

We present a conditional text generation framework that posits sentential expressions of possible causes and effects. This framework depends on two novel resources we develop in th…

cs.CL20212.5k cited

LoRA: Low-Rank Adaptation of Large Language Models

Edward J. Hu, Yelong Shen, Phillip Wallis +5

An important paradigm of natural language processing consists of large-scale pre-training on general domain data and adaptation to particular tasks or domains. As we pre-train larg…

cs.CL2020

Iterative Paraphrastic Augmentation with Discriminative Span Alignment

Ryan Culkin, J. Edward Hu, Elias Stengel-Eskin +2

We introduce a novel paraphrastic augmentation strategy based on sentence-level lexically constrained paraphrasing and discriminative span alignment. Our approach allows for the la…

cs.LG2020

Randomized Smoothing of All Shapes and Sizes

Greg Yang, Tony Duan, J. Edward Hu +3

Randomized smoothing is the current state-of-the-art defense with provable robustness against adversarial attacks. Many works have devised new randomized smoothing schemes…

cs.CL20191 cited

ParaBank: Monolingual Bitext Generation and Sentential Paraphrasing via Lexically-constrained Neural Machine Translation

J. Edward Hu, Rachel Rudinger, Matt Post +1

We present ParaBank, a large-scale English paraphrase dataset that surpasses prior work in both quantity and quality. Following the approach of ParaNMT, we train a Czech-English ne…