3 papers
cs.LG2026
Multi-Headed Transformer Architectures as Time-dependent Wasserstein Gradient Flows
Alex Massucco, Leonardo Del Grande, Marcello Carioni +2
In recent years, transformer architectures have revolutionized the field of language processing, opening the door to previously unforeseen possibilities. However, from a theoretica…
math.OC2026
A Dual Certificate Approach to Sparsity in Infinite-Width Shallow Neural Networks
Leonardo Del Grande, Christoph Brune, Marcello Carioni
In this paper, we study total variation (TV)-regularized training of infinite-width shallow ReLU neural networks, formulated as a convex optimization problem over measures on the u…
math.OC2024
Exact Sparse Representation Recovery in Signal Demixing and Group BLASSO
Marcello Carioni, Leonardo Del Grande
In this short article we present the theory of sparse representations recovery in convex regularized optimization problems introduced in (Carioni and Del Grande, arXiv:2311.08072,…