4 papers
Towards Worst-Case Guarantees with Scale-Aware Interpretability
Lauren Greenspan, David Berman, Aryeh Brill +9
Neural networks organize information according to the hierarchical, multi-scale structure of natural data. Methods to interpret model internals should be similarly scale-aware, exp…
A path to natural language through tokenisation and transformers
David S. Berman, Alexander G. Stapleton
Natural languages exhibit striking regularities in their statistical structure, including notably the emergence of Zipf's and Heaps' laws. Despite this, it remains broadly unclear…
Grokking vs. Learning: Same Features, Different Encodings
Dmitry Manning-Coe, Jacopo Gliozzi, Alexander G. Stapleton +4
Grokking typically achieves similar loss to ordinary, "steady", learning. We ask whether these different learning paths - grokking versus ordinary training - lead to fundamental di…
AInstein: Numerical Einstein Metrics via Machine Learning
Edward Hirst, Tancredi Schettini Gherardini, Alexander G. Stapleton
A new semi-supervised machine learning package is introduced which successfully solves the Euclidean vacuum Einstein equations with a cosmological constant, without any symmetry as…