2 papers
cs.LG2026
Transformers learn factored representations
Adam Shai, Loren Amdahl-Culleton, Casper L. Christensen +6
Transformers pretrained via next token prediction learn to factor their world into parts, representing these factors in orthogonal subspaces of the residual stream. We formalize tw…
cs.LG2025
Next-token pretraining implies in-context learning
Paul M. Riechers, Henry R. Bigelow, Eric A. Alt +1
We argue that in-context learning (ICL) predictably arises from standard self-supervised next-token pretraining, rather than being an exotic emergent property. This work establishe…