2 papers
cs.LG2026
Transformers learn factored representations
Adam Shai, Loren Amdahl-Culleton, Casper L. Christensen +6
Transformers pretrained via next token prediction learn to factor their world into parts, representing these factors in orthogonal subspaces of the residual stream. We formalize tw…
cond-mat.stat-mech2025
Large Interconnected Thermodynamic Systems Nearly Minimize Entropy Production
Kyle J. Ray, Alexander B. Boyd
Many have speculated whether nonequilibrium systems obey principles of maximum or minimum entropy production. In this work, we use stochastic thermodynamics to derive the condition…