papers
Publications (3)
cs.LG2025
Constrained belief updates explain geometric structures in transformer representations
Mateusz Piotrowski, Paul M. Riechers, Daniel Filan +1
What computational structures emerge in transformers trained on next-token prediction? In this work, we provide evidence that transformers implement constrained Bayesian belief upd…
cs.LG2025
Neural networks leverage nominally quantum and post-quantum representations
Paul M. Riechers, Thomas J. Elliott, Adam S. Shai
We show that deep neural networks, including transformers and RNNs, pretrained as usual on next-token prediction, intrinsically discover and represent beliefs over 'quantum' and 'p…
cs.LG2025
Transformers represent belief state geometry in their residual stream
Adam S. Shai, Sarah E. Marzen, Lucas Teixeira +2
What computational structure are we building into large language models when we train them on next-token prediction? Here, we present evidence that this structure is given by the m…