activity
20242026
collaborators

16 papers

cs.LG2026

Singular Learning and Occam's Razor in Deep Monomial Networks

Kathlén Kohn, Giovanni Luca Marchetti, Farhan Shabir +2

In the optimization of neural networks, gradient dynamics are influenced by critical points that arise from the model's architecture. These critical points occur where the Jacobian…

cs.LG2026

On the Geometry and Optimization of Polynomial Convolutional Networks

Vahid Shahverdi, Giovanni Luca Marchetti, Kathlén Kohn

We study convolutional neural networks with monomial activation functions. Specifically, we prove that their parameterization map is regular and is an isomorphism almost everywhere…

cs.LG2026

Learning on a Razor's Edge: Identifiability and Singularity of Polynomial Neural Networks

Vahid Shahverdi, Giovanni Luca Marchetti, Kathlén Kohn

We study function spaces parametrized by neural networks, referred to as neuromanifolds. Specifically, we focus on deep Multi-Layer Perceptrons (MLPs) and Convolutional Neural Netw…

cs.LG2026

Geometry of Lightning Self-Attention: Identifiability and Dimension

Nathan W. Henry, Giovanni Luca Marchetti, Kathlén Kohn

We consider function spaces defined by self-attention networks without normalization, and theoretically analyze their geometry. Since these networks are polynomial, we rely on tool…

cs.LG2026

Identifiable Equivariant Networks are Layerwise Equivariant

Vahid Shahverdi, Giovanni Luca Marchetti, Georg Bökman +1

We investigate the relation between end-to-end equivariance and layerwise equivariance in deep neural networks. We prove the following: For a network whose end-to-end function is e…

cs.LG2026

The Riemannian Geometry Associated to Gradient Flows of Linear Convolutional Networks

El Mehdi Achour, Kathlén Kohn, Holger Rauhut

We study geometric properties of the gradient flow for learning deep linear convolutional networks. For linear fully connected networks, it has been shown recently that the corresp…