4 papers
Multi-Headed Transformer Architectures as Time-dependent Wasserstein Gradient Flows
Alex Massucco, Leonardo Del Grande, Marcello Carioni +2
In recent years, transformer architectures have revolutionized the field of language processing, opening the door to previously unforeseen possibilities. However, from a theoretica…
Muon is Not That Special: Random or Inverted Spectra Work Just as Well
Zakhar Shumaylov, Nathaël Da Costa, Peter Zaika +6
The recent empirical success of the Muon optimizer has renewed interest in non-Euclidean optimization, typically justified by similarities with second-order methods, and linear min…
Deep Network Trainability via Persistent Subspace Orthogonality
Alex Massucco, Davide Murari, Carola-Bibiane Schönlieb
Training neural networks via backpropagation is often hindered by vanishing or exploding gradients. In this work, we design architectures that mitigate these issues by analyzing an…
Finite-difference least square methods for solving Hamilton-Jacobi equations using neural networks
Carlos Esteve-Yagüe, Richard Tsai, Alex Massucco
We present a simple algorithm to approximate the viscosity solution of Hamilton-Jacobi (HJ) equations by means of an artificial deep neural network. The algorithm uses a stochastic…