26 citations · 76 across the 8 of their papers we have counts for
9 papers
What Language Model to Train if You Have One Million GPU Hours?
Teven Le Scao, Thomas Wang, Daniel Hesslow +16
The crystallization of modeling methods around the Transformer architecture has been a boon for practitioners. Simple, well-motivated architectural variations can transfer across t…
Scaling Laws Beyond Backpropagation
Matthew J. Filipovich, Alessandro Cappelli, Daniel Hesslow +1
Alternatives to backpropagation have long been studied to better understand how biological brains may learn. Recently, they have also garnered interest as a way to train neural net…
What Language Model Architecture and Pretraining Objective Work Best for Zero-Shot Generalization?
Thomas Wang, Adam Roberts, Daniel Hesslow +5
Large pretrained Transformer language models have been shown to exhibit zero-shot generalization, i.e. they can perform a wide variety of tasks that they were not explicitly traine…
Is the Number of Trainable Parameters All That Actually Matters?
Amélie Chatelain, Amine Djeghri, Daniel Hesslow +2
Recent work has identified simple empirical scaling laws for language models, linking compute budget, dataset size, model size, and autoregressive modeling loss. The validity of th…
ROPUST: Improving Robustness through Fine-tuning with Photonic Processors and Synthetic Gradients
Alessandro Cappelli, Julien Launay, Laurent Meunier +2
Robustness to adversarial attacks is typically obtained through expensive adversarial training with Projected Gradient Descent. Here we introduce ROPUST, a remarkably simple and ef…
Hardware Beyond Backpropagation: a Photonic Co-Processor for Direct Feedback Alignment
Julien Launay, Iacopo Poli, Kilian Müller +5
The scaling hypothesis motivates the expansion of models past trillions of parameters as a path towards better performance. Recent significant developments, such as GPT-3, have bee…