4 papers
Tunneling the Loss Landscape: Bypassing Memorization with Monte Carlo Parameter Swapping
Lai Shun Chan, Xiaotian Zhang, Yue Shang +2
Grokking is a striking phenomenon in neural network training, where a model can undergo a prolonged period of pure memorization before abrupt generalization. While previous works h…
The Grokked Illusion: True Equilibrium Mitigates Catastrophic Forgetting
Xiaotian Zhang, Lai Shun Chan, Yue Shang +2
While neural networks are typically evaluated by their training and test performance, these metrics do not reveal how robust a learned representation is. Recent studies have shown…
Is Grokking a Computational Glass Relaxation?
Xiaotian Zhang, Yue Shang, Entao Yang +1
Understanding neural network's (NN) generalizability remains a central question in deep learning research. The special phenomenon of grokking, where NNs abruptly generalize long af…
High-entropy Advantage in Neural Networks' Generalizability
Entao Yang, Xiaotian Zhang, Yue Shang +1
One of the central challenges in modern machine learning is understanding how neural networks generalize knowledge learned from training data to unseen test data. While numerous em…