1 citations · 1 across the 3 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Is Grokking a Loss of Normal Hyperbolicity of the Interpolation Manifold?
Suvinava Basak
A recent line of work recasts the post-memorization phase of grokking as constrained optimization: once a network interpolates the training set, weight decay drives a slow drift al…
cs.LG2026
ZC-Swish: Stabilizing Deep BN-Free Networks for Edge and Micro-Batch Applications
Suvinava Basak
Batch Normalization (BN) is a cornerstone of deep learning, yet it fundamentally breaks down in micro-batch regimes (e.g., 3D medical imaging) and non-IID Federated Learning. Remov…