3 papers
cs.LG2026
Late-Stage Generalization Collapse in Grokking: Detecting anti-grokking with Weightwatcher
Hari K Prakash, Charles H Martin
\emph{Memorization} in neural networks lacks a precise operational definition and is often inferred from the grokking regime, where training accuracy saturates while test accuracy…
cs.LG2025
SETOL: A Semi-Empirical Theory of (Deep) Learning
Charles H Martin, Christopher Hinrichs
We present a SemiEmpirical Theory of Learning (SETOL) that explains the remarkable performance of State-Of-The-Art (SOTA) Neural Networks (NNs). We provide a formal explanation of…
cs.LG2025
Grokking and Generalization Collapse: Insights from \texttt{HTSR} theory
Hari K. Prakash, Charles H. Martin
We study the well-known grokking phenomena in neural networks (NNs) using a 3-layer MLP trained on 1 k-sample subset of MNIST, with and without weight decay, and discover a novel t…