6 papers · 1 filter
Transformers learn factored representations
Adam Shai, Loren Amdahl-Culleton, Casper L. Christensen +6
Transformers pretrained via next token prediction learn to factor their world into parts, representing these factors in orthogonal subspaces of the residual stream. We formalize tw…
Rank-1 LoRAs Encode Interpretable Reasoning Signals
Jake Ward, Paul Riechers, Adam Shai
Reasoning models leverage inference-time compute to significantly enhance the performance of language models on difficult logical tasks, and have become a dominating paradigm in fr…
Constrained belief updates explain geometric structures in transformer representations
Mateusz Piotrowski, Paul M. Riechers, Daniel Filan +1
What computational structures emerge in transformers trained on next-token prediction? In this work, we provide evidence that transformers implement constrained Bayesian belief upd…
Neural networks leverage nominally quantum and post-quantum representations
Paul M. Riechers, Thomas J. Elliott, Adam S. Shai
We show that deep neural networks, including transformers and RNNs, pretrained as usual on next-token prediction, intrinsically discover and represent beliefs over 'quantum' and 'p…
Next-token pretraining implies in-context learning
Paul M. Riechers, Henry R. Bigelow, Eric A. Alt +1
We argue that in-context learning (ICL) predictably arises from standard self-supervised next-token pretraining, rather than being an exotic emergent property. This work establishe…
Transformers represent belief state geometry in their residual stream
Adam S. Shai, Sarah E. Marzen, Lucas Teixeira +2
What computational structure are we building into large language models when we train them on next-token prediction? Here, we present evidence that this structure is given by the m…