16 papers
When Do Attention Circuits Form? Developmental Trajectories of Capability and Attention-Sink Emergence Across Three 1B-ClassArchitectures
Yongzhong Xu
We track the developmental trajectory of attention-head circuit formation across three 1B-class language models spanning two architecture families (dense transformer, mixture-of-ex…
Closure-Validated Circuit Discovery in Attention Heads: Co-activation Proposes, Ablation Disposes
Yongzhong Xu
Interpretability increasingly treats groups of components, not individual units, as the basic object, and proposes to find them by clustering co-activation statistics. We ask wheth…
Pattern Selectivity is Not Task-Causal Structure: A Cross-Architecture Mechanistic Study of Composed-Task Circuits in 1B-Class Language Models
Yongzhong Xu
We test whether a single screen-and-ablate recipe -- identify attention-head circuits by task-pattern selectivity, then verify by causal ablation against a matched-random null -- p…
Spectral Probe-Circuits: A Three-Step Recipe for Identifying Attention-Head Circuits in Pretrained Transformers
Yongzhong Xu
We present a three-step recipe for identifying attention-head circuits in pretrained transformers. A per-head spectral signal -- the time-integrated participation ratio of each hea…
Spectral Edge Dynamics: An Analytical-Empirical Study of Phase Transitions in Neural Network Training
Yongzhong Xu
We develop the spectral edge analysis: phase transitions in neural network training -- grokking, capability gains, loss plateaus -- are controlled by the spectral gap of the rollin…
Feature Repulsion and Spectral Lock-in: An Empirical Study of Two-Layer Network Grokking
Yongzhong Xu
Tian (2025) proves a repulsion theorem (Theorem 6) for the matrix during the interactive feature-learning stage of grokking: s…