1.5k citations · 1.9k across the 8 of their papers we have counts for
19 papers
Language Models are Few-shot Multilingual Learners
Genta Indra Winata, Andrea Madotto, Zhaojiang Lin +3
General-purpose language models have demonstrated impressive capabilities, performing on par with state-of-the-art approaches on a range of downstream natural language processing (…
When does loss-based prioritization fail?
Niel Teng Hu, Xinyu Hu, Rosanne Liu +2
Not all examples are created equal, but standard deep neural network training protocols treat each training point uniformly. Each example is propagated forward and backward through…
Supermasks in Superposition
Mitchell Wortsman, Vivek Ramanujan, Rosanne Liu +4
We present the Supermasks in Superposition (SupSup) model, capable of sequentially learning thousands of tasks without catastrophic forgetting. Our approach uses a randomly initial…
Estimating Q(s,s') with Deep Deterministic Dynamics Gradients
Ashley D. Edwards, Himanshu Sahni, Rosanne Liu +7
In this paper, we introduce a novel form of value function, , that expresses the utility of transitioning from a state to a neighboring state and then acting opt…
Plug and Play Language Models: A Simple Approach to Controlled Text Generation
Sumanth Dathathri, Andrea Madotto, Janice Lan +5
Large transformer-based language models (LMs) trained on huge text corpora have shown unparalleled generation capabilities. However, controlling attributes of the generated languag…
First-Order Preconditioning via Hypergradient Descent
Ted Moskovitz, Rui Wang, Janice Lan +4
Standard gradient descent methods are susceptible to a range of issues that can impede training, such as high correlations and different scaling in parameter space.These difficulti…