1 citations · 1 across the 5 of their papers we have counts for
5 papers · 1 filter
Universal Length Generalization with Turing Programs
Kaiying Hou, David Brandfonbrener, Sham Kakade +2
Length generalization refers to the ability to extrapolate from short training sequences to long test sequences and is a challenge for current large language models. While prior wo…
A New Perspective on Shampoo's Preconditioner
Depen Morwani, Itai Shapira, Nikhil Vyas +3
Shampoo, a second-order optimization algorithm which uses a Kronecker product preconditioner, has recently garnered increasing attention from the machine learning community. The pr…
The Evolution of Statistical Induction Heads: In-Context Learning Markov Chains
Benjamin L. Edelman, Ezra Edelman, Surbhi Goel +2
Large language models have the ability to generate text that mimics patterns in their inputs. We introduce a simple Markov Chain sequence modeling task in order to study how this i…
Pareto Frontiers in Neural Feature Learning: Data, Compute, Width, and Luck
Benjamin L. Edelman, Surbhi Goel, Sham Kakade +2
In modern deep learning, algorithmic choices (such as width, depth, and learning rate) are known to modulate nuanced resource tradeoffs. This work investigates how these complexiti…
Corgi^2: A Hybrid Offline-Online Approach To Storage-Aware Data Shuffling For SGD
Etay Livne, Gal Kaplun, Eran Malach +1
When using Stochastic Gradient Descent (SGD) for training machine learning models, it is often crucial to provide the model with examples sampled at random from the dataset. Howeve…