machine learning

Gradient Flow Dynamics and Implicit Bias of Diagonal Linear Networks under Infinitesimal Initialization

arXiv:2607.12332

summary

The paper analyzes how diagonal linear networks trained with infinitesimal initialization evolve under gradient flow, showing they follow a specific algorithm that converges to a modified \(l_1\) norm solution, thereby revealing the networks' implicit bias and the role of a structural invariant manifold.

Abstract

We study the gradient flow dynamics of diagonal linear networks for regression tasks under infinitesimal initialization. Extending Theorem 1 from Pesme & Flammarion (2023), we generalize the analysis to both deep diagonal linear networks and a broader class of two-layer diagonal linear networks (as defined in Definition 4.1). Specifically, we demonstrate that the training trajectories of these models can be equivalently characterized by the proposed Algorithm 1. We further prove that this algorithm converges to the solution of a modified norm minimization problem. As a result, we establish that the implicit bias of both network architectures corresponds to a modified norm in the regime of infinitesimal initialization. Additionally, we provide insights into the underlying mechanisms governing these dynamics by identifying the Structural Invariant Manifold (SIM) (Zhao et al., 2026) as the key geometric structure that shapes the learning process.

Topics & keywords

#gradient flow#diagonal linear networks#implicit bias#infinitesimal initialization#l1 norm minimizationgradient flow dynamicsdiagonal linear networkmodified l1 normstructural invariant manifoldalgorithmic equivalence
Gradient Flow Dynamics and Implicit Bias of Diagonal Linear Networks under Infinitesimal Initialization · wovepaper