571 citations · 1.2k across the 12 of their papers we have counts for
4 papers · 1 filter
Benchmarking Neural Network Training Algorithms
George E. Dahl, Frank Schneider, Zachary Nado +22
Training algorithms, broadly construed, are an essential part of every deep learning pipeline. Training algorithm improvements that speed up training across a wide variety of workl…
Unifying Grokking and Double Descent
Xander Davies, Lauro Langosco, David Krueger
A principled understanding of generalization in deep learning may require unifying disparate observations under a single conceptual framework. Previous work has studied \emph{grokk…
Scaling Vision Transformers to 22 Billion Parameters
Mostafa Dehghani, Josip Djolonga, Basil Mustafa +39
The scaling of Transformers has driven breakthrough capabilities for language models. At present, the largest large language models (LLMs) contain upwards of 100B parameters. Visio…
Improving Training Stability for Multitask Ranking Models in Recommender Systems
Jiaxi Tang, Yoel Drori, Daryl Chang +6
Recommender systems play an important role in many content platforms. While most recommendation research is dedicated to designing better models to improve user experience, we foun…