5 papers · 1 filter
Finite-Time Analysis of Gradient Descent for Shallow Transformers
Enes Arda, Semih Cayci, Atilla Eryilmaz
Understanding why Transformers perform so well remains challenging due to their non-convex optimization landscape. In this work, we analyze a shallow Transformer with independe…
Convergence of Stochastic Gradient Langevin Dynamics in the Lazy Training Regime
Noah Oberweis, Semih Cayci
Continuous-time models provide important insights into the training dynamics of optimization algorithms in deep learning. In this work, we establish a non-asymptotic convergence an…
Non-Asymptotic Optimization and Generalization Bounds for Stochastic Gauss-Newton in Overparameterized Models
Semih Cayci
An important question in deep learning is how higher-order optimization methods affect generalization. In this work, we analyze a stochastic Gauss-Newton (SGN) method with Levenber…
Convergence of Gradient Descent for Recurrent Neural Networks: A Nonasymptotic Analysis
Semih Cayci, Atilla Eryilmaz
We analyze recurrent neural networks with diagonal hidden-to-hidden weight matrices, trained with gradient descent in the supervised learning setting, and prove that gradient desce…
Linear Convergence of Entropy-Regularized Natural Policy Gradient with Linear Function Approximation
Semih Cayci, Niao He, R. Srikant
Natural policy gradient (NPG) methods with entropy regularization achieve impressive empirical success in reinforcement learning problems with large state-action spaces. However, t…