1 paper
Adam Shai, Loren Amdahl-Culleton, Casper L. Christensen +6
Transformers pretrained via next token prediction learn to factor their world into parts, representing these factors in orthogonal subspaces of the residual stream. We formalize tw…