most citedThe Evolution of Statistical Induction Heads: In-Context Learning Markov Chains

1 citations · 1 across the 5 of their papers we have counts for

collaborators
Showing cs.LGShow all

5 papers · 1 filter

cs.LG2024

Universal Length Generalization with Turing Programs

Kaiying Hou, David Brandfonbrener, Sham Kakade +2

Length generalization refers to the ability to extrapolate from short training sequences to long test sequences and is a challenge for current large language models. While prior wo…

cs.LG2024

A New Perspective on Shampoo's Preconditioner

Depen Morwani, Itai Shapira, Nikhil Vyas +3

Shampoo, a second-order optimization algorithm which uses a Kronecker product preconditioner, has recently garnered increasing attention from the machine learning community. The pr…

cs.LG20241 cited

The Evolution of Statistical Induction Heads: In-Context Learning Markov Chains

Benjamin L. Edelman, Ezra Edelman, Surbhi Goel +2

Large language models have the ability to generate text that mimics patterns in their inputs. We introduce a simple Markov Chain sequence modeling task in order to study how this i…

cs.LG2023

Pareto Frontiers in Neural Feature Learning: Data, Compute, Width, and Luck

Benjamin L. Edelman, Surbhi Goel, Sham Kakade +2

In modern deep learning, algorithmic choices (such as width, depth, and learning rate) are known to modulate nuanced resource tradeoffs. This work investigates how these complexiti…

cs.LG2023

Corgi^2: A Hybrid Offline-Online Approach To Storage-Aware Data Shuffling For SGD

Etay Livne, Gal Kaplun, Eran Malach +1

When using Stochastic Gradient Descent (SGD) for training machine learning models, it is often crucial to provide the model with examples sampled at random from the dataset. Howeve…