11 papers
On the Diverse Dynamical Behaviors Arising in Deep Linear Transformers
Sixu Li, Thomas Jacob Maranzatto, Jan Peszek +5
We study the inference-time behavior of deep linear encoder-only transformers through the lens of interacting particle systems. In this perspective, tokens are modeled as particles…
Exploiting Structure with Anisotropic Consensus-Based Optimization
Sabrina Bonandin, Konstantin Riedl, Sara Veneruso
Anisotropic consensus-based optimization (CBO), a multi-agent metaheuristic derivative-free optimization method, which reliably finds global minima of nonsmooth and nonconvex objec…
Convergence Analysis of Newton's Method for Neural Networks in the Overparameterized Limit
Konstantin Riedl, Konstantinos Spiliopoulos, Justin Sirignano
A convergence analysis is developed for the regularized Newton method for training neural networks (NNs) in the overparameterized limit. As the number of hidden units tends to infi…
Convergence of Consensus-Based Particle Methods for Nonconvex Bi-Level Optimization
Yutong Chao, Xudong Sun, Konstantin Riedl +2
In this paper, we study a consensus-based optimization method for nonconvex bi-level optimization, where the objective is to minimize an upper-level function over the set of global…
Quantifying Concentration Phenomena of Mean-Field Transformers in the Low-Temperature Regime
Albert Alcalde, Leon Bungert, Konstantin Riedl +1
Transformers with self-attention modules as their core components have become an integral architecture in modern large language and foundation models. In this paper, we study the e…
Gradient is All You Need? How Consensus-Based Optimization can be Interpreted as a Stochastic Relaxation of Gradient Descent
Konstantin Riedl, Timo Klock, Carina Geldhauser +1
In this paper, we provide a novel analytical perspective on the theoretical understanding of gradient-based learning algorithms by interpreting consensus-based optimization (CBO),…