2.5k citations · 2.6k across the 4 of their papers we have counts for
7 papers
Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer
Greg Yang, Edward J. Hu, Igor Babuschkin +7
Hyperparameter (HP) tuning in deep learning is an expensive process, prohibitively so for neural networks (NNs) with billions of parameters. We show that, in the recently discovere…
Guided Generation of Cause and Effect
Zhongyang Li, Xiao Ding, Ting Liu +2
We present a conditional text generation framework that posits sentential expressions of possible causes and effects. This framework depends on two novel resources we develop in th…
LoRA: Low-Rank Adaptation of Large Language Models
Edward J. Hu, Yelong Shen, Phillip Wallis +5
An important paradigm of natural language processing consists of large-scale pre-training on general domain data and adaptation to particular tasks or domains. As we pre-train larg…
Iterative Paraphrastic Augmentation with Discriminative Span Alignment
Ryan Culkin, J. Edward Hu, Elias Stengel-Eskin +2
We introduce a novel paraphrastic augmentation strategy based on sentence-level lexically constrained paraphrasing and discriminative span alignment. Our approach allows for the la…
Randomized Smoothing of All Shapes and Sizes
Greg Yang, Tony Duan, J. Edward Hu +3
Randomized smoothing is the current state-of-the-art defense with provable robustness against adversarial attacks. Many works have devised new randomized smoothing schemes…
ParaBank: Monolingual Bitext Generation and Sentential Paraphrasing via Lexically-constrained Neural Machine Translation
J. Edward Hu, Rachel Rudinger, Matt Post +1
We present ParaBank, a large-scale English paraphrase dataset that surpasses prior work in both quantity and quality. Following the approach of ParaNMT, we train a Czech-English ne…