3 papers
cs.LG2026
Detecting overfitting in Neural Networks during long-horizon grokking using Random Matrix Theory
Hari K. Prakash, Charles H Martin
Training Neural Networks (NNs) without overfitting is difficult; detecting that overfitting is difficult as well. We present a novel Random Matrix Theory method that detects the on…
cs.LG2026
Late-Stage Generalization Collapse in Grokking: Detecting anti-grokking with Weightwatcher
Hari K Prakash, Charles H Martin
\emph{Memorization} in neural networks lacks a precise operational definition and is often inferred from the grokking regime, where training accuracy saturates while test accuracy…
cs.LG2025
Grokking and Generalization Collapse: Insights from \texttt{HTSR} theory
Hari K. Prakash, Charles H. Martin
We study the well-known grokking phenomena in neural networks (NNs) using a 3-layer MLP trained on 1 k-sample subset of MNIST, with and without weight decay, and discover a novel t…