157 citations · 233 across the 10 of their papers we have counts for
3 papers · 1 filter
Is the Number of Trainable Parameters All That Actually Matters?
Amélie Chatelain, Amine Djeghri, Daniel Hesslow +2
Recent work has identified simple empirical scaling laws for language models, linking compute budget, dataset size, model size, and autoregressive modeling loss. The validity of th…
Direct Feedback Alignment Scales to Modern Deep Learning Tasks and Architectures
Julien Launay, Iacopo Poli, François Boniface +1
Despite being the workhorse of deep learning, the backpropagation algorithm is no panacea. It enforces sequential layer updates, thus preventing efficient parallelization of the tr…
Principled Training of Neural Networks with Direct Feedback Alignment
Julien Launay, Iacopo Poli, Florent Krzakala
The backpropagation algorithm has long been the canonical training method for neural networks. Modern paradigms are implicitly optimized for it, and numerous guidelines exist to en…