2 papers
cond-mat.dis-nn2026
Tunneling the Loss Landscape: Bypassing Memorization with Monte Carlo Parameter Swapping
Lai Shun Chan, Xiaotian Zhang, Yue Shang +2
Grokking is a striking phenomenon in neural network training, where a model can undergo a prolonged period of pure memorization before abrupt generalization. While previous works h…
cs.LG2026
The Grokked Illusion: True Equilibrium Mitigates Catastrophic Forgetting
Xiaotian Zhang, Lai Shun Chan, Yue Shang +2
While neural networks are typically evaluated by their training and test performance, these metrics do not reveal how robust a learned representation is. Recent studies have shown…