2 papers
cs.LG2026
The Global Empirical NTK: Self-Referential Bias and Dimensionality of Gradient Descent Learning
James Hazelden, Laura Driscoll, Eli Shlizerman +1
In training a neural network with gradient descent (GD), each iteration induces a linear operator that governs first-order updates to a model's internal state variables. We define…
cs.LG2025
KPFlow: An Operator Perspective on Dynamic Collapse Under Gradient Descent Training of Recurrent Networks
James Hazelden, Laura Driscoll, Eli Shlizerman +1
Gradient Descent (GD) and its variants are the primary tool for enabling efficient training of recurrent dynamical systems such as Recurrent Neural Networks (RNNs), Neural ODEs and…