3 citations · 5 across the 10 of their papers we have counts for
14 papers · 1 filter
Singular Learning and Occam's Razor in Deep Monomial Networks
Kathlén Kohn, Giovanni Luca Marchetti, Farhan Shabir +2
In the optimization of neural networks, gradient dynamics are influenced by critical points that arise from the model's architecture. These critical points occur where the Jacobian…
On the Geometry and Optimization of Polynomial Convolutional Networks
Vahid Shahverdi, Giovanni Luca Marchetti, Kathlén Kohn
We study convolutional neural networks with monomial activation functions. Specifically, we prove that their parameterization map is regular and is an isomorphism almost everywhere…
Equivariant Representation Learning via Class-Pose Decomposition
Giovanni Luca Marchetti, Gustaf Tegnér, Anastasiia Varava +1
We introduce a general method for learning representations that are equivariant to symmetries of data. Our central idea is to decompose the latent space into an invariant factor an…
Learning on a Razor's Edge: Identifiability and Singularity of Polynomial Neural Networks
Vahid Shahverdi, Giovanni Luca Marchetti, Kathlén Kohn
We study function spaces parametrized by neural networks, referred to as neuromanifolds. Specifically, we focus on deep Multi-Layer Perceptrons (MLPs) and Convolutional Neural Netw…
Geometry of Lightning Self-Attention: Identifiability and Dimension
Nathan W. Henry, Giovanni Luca Marchetti, Kathlén Kohn
We consider function spaces defined by self-attention networks without normalization, and theoretically analyze their geometry. Since these networks are polynomial, we rely on tool…
Sequential Group Composition: A Window into the Mechanics of Deep Learning
Giovanni Luca Marchetti, Daniel Kunin, Adele Myers +2
How do neural networks trained over sequences acquire the ability to perform structured operations, such as arithmetic, geometric, and algorithmic computation? To gain insight into…