141 citations · 141 across the 1 of their papers we have counts for
4 papers
Plug and Play Language Models: A Simple Approach to Controlled Text Generation
Sumanth Dathathri, Andrea Madotto, Janice Lan +5
Large transformer-based language models (LMs) trained on huge text corpora have shown unparalleled generation capabilities. However, controlling attributes of the generated languag…
First-Order Preconditioning via Hypergradient Descent
Ted Moskovitz, Rui Wang, Janice Lan +4
Standard gradient descent methods are susceptible to a range of issues that can impede training, such as high correlations and different scaling in parameter space.These difficulti…
LCA: Loss Change Allocation for Neural Network Training
Janice Lan, Rosanne Liu, Hattie Zhou +1
Neural networks enjoy widespread use, but many aspects of their training, representation, and operation are poorly understood. In particular, our view into the training process is…
Deconstructing Lottery Tickets: Zeros, Signs, and the Supermask
Hattie Zhou, Janice Lan, Rosanne Liu +1
The recent "Lottery Ticket Hypothesis" paper by Frankle & Carbin showed that a simple approach to creating sparse networks (keeping the large weights) results in models that are tr…