7 papers
Beyond Coordinate Gauge: An Audited Protocol for Detecting Donor-Specific Functional Fingerprints after Neural Collapse
Truong Xuan Khanh, Phan Thanh Duc
Independently trained neural networks have no shared neuron-index reference frame, so comparing them requires accounting for coordinate freedom. Neural Collapse sharpens this probl…
At-Grok Is Not Converged:A Measurement-Validity Audit for Grokking Representation Metrics
Truong Xuan Khanh
On modular arithmetic, a network's embedding keeps compressing for tens of thousands of steps after it has already generalized. Reading effective rank at the grokking transition ov…
Cross-Trajectory Chimera Interventions Reveal Dissociable Roles of Weight Magnitude and Direction in Grokking
Truong Xuan Khanh
Which properties of a partially trained network are causally portable to a different, independently trained network? Single-trajectory interventions show necessity within one run,…
What Does the Weight Norm Control in Grokking? Logit-Scale Mediation under Cross-Entropy
Truong Xuan Khanh
Grokking, the delayed jump from memorization to generalization, is usually tied to the weight norm: a smaller norm generalizes sooner. We ask what the norm actually controls. Holdi…
The Weight Norm Sets the Grokking Timescale: A Causal Delay Law
Truong Xuan Khanh, Doan Hoang Viet, Luu Duc Trung +1
Grokking is the delayed onset of generalization in neural networks, arising long after they fit the training data. Whether the weight norm causes this delay is disputed: some studi…
Norm-Hierarchy Transitions in Representation Learning: When and Why Neural Networks Abandon Shortcuts
Truong Xuan Khanh, Truong Quynh Hoa
Neural networks often rely on spurious shortcuts for many epochs before discovering structured representations. However, the mechanism governing when this transition occurs and whe…