3 papers
hep-th2026
Towards Worst-Case Guarantees with Scale-Aware Interpretability
Lauren Greenspan, David Berman, Aryeh Brill +9
Neural networks organize information according to the hierarchical, multi-scale structure of natural data. Methods to interpret model internals should be similarly scale-aware, exp…
cs.LG2025
Exact Learning Dynamics of In-Context Learning in Linear Transformers and Its Application to Non-Linear Transformers
Nischal Mainali, Lucas Teixeira
Transformer models exhibit remarkable in-context learning (ICL), adapting to novel tasks from examples within their context, yet the underlying mechanisms remain largely mysterious…
cs.LG2025
Transformers represent belief state geometry in their residual stream
Adam S. Shai, Sarah E. Marzen, Lucas Teixeira +2
What computational structure are we building into large language models when we train them on next-token prediction? Here, we present evidence that this structure is given by the m…