Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
A simple connection from loss flatness to compressed neural representations
Shirui Chen, Stefano Recanatesi, Eric Shea-Brown
Despite extensive study, the significance of sharpness -- the trace of the loss Hessian at local minima -- remains unclear. We investigate an alternative perspective: how sharpness…
cs.LG2025
KPFlow: An Operator Perspective on Dynamic Collapse Under Gradient Descent Training of Recurrent Networks
James Hazelden, Laura Driscoll, Eli Shlizerman +1
Gradient Descent (GD) and its variants are the primary tool for enabling efficient training of recurrent dynamical systems such as Recurrent Neural Networks (RNNs), Neural ODEs and…