15 papers
Beyond Coordinate Gauge: An Audited Protocol for Detecting Donor-Specific Functional Fingerprints after Neural Collapse
Truong Xuan Khanh, Phan Thanh Duc
Independently trained neural networks have no shared neuron-index reference frame, so comparing them requires accounting for coordinate freedom. Neural Collapse sharpens this probl…
At-Grok Is Not Converged:A Measurement-Validity Audit for Grokking Representation Metrics
Truong Xuan Khanh
On modular arithmetic, a network's embedding keeps compressing for tens of thousands of steps after it has already generalized. Reading effective rank at the grokking transition ov…
Cross-Trajectory Chimera Interventions Reveal Dissociable Roles of Weight Magnitude and Direction in Grokking
Truong Xuan Khanh
Which properties of a partially trained network are causally portable to a different, independently trained network? Single-trajectory interventions show necessity within one run,…
What Does the Weight Norm Control in Grokking? Logit-Scale Mediation under Cross-Entropy
Truong Xuan Khanh
Grokking, the delayed jump from memorization to generalization, is usually tied to the weight norm: a smaller norm generalizes sooner. We ask what the norm actually controls. Holdi…
The Weight Norm Sets the Grokking Timescale: A Causal Delay Law
Truong Xuan Khanh, Doan Hoang Viet, Luu Duc Trung +1
Grokking is the delayed onset of generalization in neural networks, arising long after they fit the training data. Whether the weight norm causes this delay is disputed: some studi…
Emergence via Phase Transitions: Mechanism Landscapes and Universal Convergence Across Complex Systems
Truong Xuan Khanh
Across machine learning, biology, and physics, independently evolving systems often converge toward strikingly similar high-level structures despite radically different microscopic…