3 papers
cs.LG2026
Towards Spectroscopy: Susceptibility Clusters in Language Models
Andrew Gordon, Garrett Baker, George Wang +3
Spectroscopy infers the internal structure of physical systems by measuring their response to perturbations. We apply this principle to neural networks: perturbing the data distrib…
cs.LG2025
SCALAR: Benchmarking SAE Interaction Sparsity in Toy LLMs
Sean P. Fillingham, Andrew Gordon, Peter Lai +3
Mechanistic interpretability aims to decompose neural networks into interpretable features and map their connecting circuits. The standard approach trains sparse autoencoders (SAEs…
cs.LG2025
Embryology of a Language Model
George Wang, Garrett Baker, Andrew Gordon +1
Understanding how language models develop their internal computational structure is a central problem in the science of deep learning. While susceptibilities, drawn from statistica…