16 papers
Singular Learning and Occam's Razor in Deep Monomial Networks
Kathlén Kohn, Giovanni Luca Marchetti, Farhan Shabir +2
In the optimization of neural networks, gradient dynamics are influenced by critical points that arise from the model's architecture. These critical points occur where the Jacobian…
On the Geometry and Optimization of Polynomial Convolutional Networks
Vahid Shahverdi, Giovanni Luca Marchetti, Kathlén Kohn
We study convolutional neural networks with monomial activation functions. Specifically, we prove that their parameterization map is regular and is an isomorphism almost everywhere…
Learning on a Razor's Edge: Identifiability and Singularity of Polynomial Neural Networks
Vahid Shahverdi, Giovanni Luca Marchetti, Kathlén Kohn
We study function spaces parametrized by neural networks, referred to as neuromanifolds. Specifically, we focus on deep Multi-Layer Perceptrons (MLPs) and Convolutional Neural Netw…
Geometry of Lightning Self-Attention: Identifiability and Dimension
Nathan W. Henry, Giovanni Luca Marchetti, Kathlén Kohn
We consider function spaces defined by self-attention networks without normalization, and theoretically analyze their geometry. Since these networks are polynomial, we rely on tool…
Identifiable Equivariant Networks are Layerwise Equivariant
Vahid Shahverdi, Giovanni Luca Marchetti, Georg Bökman +1
We investigate the relation between end-to-end equivariance and layerwise equivariance in deep neural networks. We prove the following: For a network whose end-to-end function is e…
The Riemannian Geometry Associated to Gradient Flows of Linear Convolutional Networks
El Mehdi Achour, Kathlén Kohn, Holger Rauhut
We study geometric properties of the gradient flow for learning deep linear convolutional networks. For linear fully connected networks, it has been shown recently that the corresp…