2 papers
cs.LG2026
Early Data Exposure Improves Robustness to Subsequent Fine-Tuning
Lawrence Feng, Gaurav R. Ghosal, Jacob Mitchell Springer +2
How can we train models whose post-trained capabilities survive subsequent fine-tuning? Rather than focusing on downstream interventions to mitigate forgetting of upstream capabili…
cs.LG2023
Grokking as Compression: A Nonlinear Complexity Perspective
Ziming Liu, Ziqian Zhong, Max Tegmark
We attribute grokking, the phenomenon where generalization is much delayed after memorization, to compression. To do so, we define linear mapping number (LMN) to measure network co…