collaborators

16 papers

cs.LG2026

When Do Attention Circuits Form? Developmental Trajectories of Capability and Attention-Sink Emergence Across Three 1B-ClassArchitectures

Yongzhong Xu

We track the developmental trajectory of attention-head circuit formation across three 1B-class language models spanning two architecture families (dense transformer, mixture-of-ex…

cs.LG2026

Closure-Validated Circuit Discovery in Attention Heads: Co-activation Proposes, Ablation Disposes

Yongzhong Xu

Interpretability increasingly treats groups of components, not individual units, as the basic object, and proposes to find them by clustering co-activation statistics. We ask wheth…

cs.LG2026

Pattern Selectivity is Not Task-Causal Structure: A Cross-Architecture Mechanistic Study of Composed-Task Circuits in 1B-Class Language Models

Yongzhong Xu

We test whether a single screen-and-ablate recipe -- identify attention-head circuits by task-pattern selectivity, then verify by causal ablation against a matched-random null -- p…

cs.LG2026

Spectral Probe-Circuits: A Three-Step Recipe for Identifying Attention-Head Circuits in Pretrained Transformers

Yongzhong Xu

We present a three-step recipe for identifying attention-head circuits in pretrained transformers. A per-head spectral signal -- the time-integrated participation ratio of each hea…

cs.LG2026

Spectral Edge Dynamics: An Analytical-Empirical Study of Phase Transitions in Neural Network Training

Yongzhong Xu

We develop the spectral edge analysis: phase transitions in neural network training -- grokking, capability gains, loss plateaus -- are controlled by the spectral gap of the rollin…

cs.LG2026

Feature Repulsion and Spectral Lock-in: An Empirical Study of Two-Layer Network Grokking

Yongzhong Xu

Tian (2025) proves a repulsion theorem (Theorem 6) for the matrix during the interactive feature-learning stage of grokking: s…