11 citations · 20 across the 9 of their papers we have counts for
9 papers · 1 filter
Finite-Time Analysis of Gradient Descent for Shallow Transformers
Enes Arda, Semih Cayci, Atilla Eryilmaz
Understanding why Transformers perform so well remains challenging due to their non-convex optimization landscape. In this work, we analyze a shallow Transformer with independe…
Non-Asymptotic Optimization and Generalization Bounds for Stochastic Gauss-Newton in Overparameterized Models
Semih Cayci
An important question in deep learning is how higher-order optimization methods affect generalization. In this work, we analyze a stochastic Gauss-Newton (SGN) method with Levenber…
Convergence of Stochastic Gradient Langevin Dynamics in the Lazy Training Regime
Noah Oberweis, Semih Cayci
Continuous-time models provide important insights into the training dynamics of optimization algorithms in deep learning. In this work, we establish a non-asymptotic convergence an…
Convergence of Gradient Descent for Recurrent Neural Networks: A Nonasymptotic Analysis
Semih Cayci, Atilla Eryilmaz
We analyze recurrent neural networks with diagonal hidden-to-hidden weight matrices, trained with gradient descent in the supervised learning setting, and prove that gradient desce…
Provably Robust Temporal Difference Learning for Heavy-Tailed Rewards
Semih Cayci, Atilla Eryilmaz
In a broad class of reinforcement learning applications, stochastic rewards have heavy-tailed distributions, which lead to infinite second-order moments for stochastic (semi)gradie…
Sample Complexity and Overparameterization Bounds for Temporal Difference Learning with Neural Network Approximation
Semih Cayci, Siddhartha Satpathi, Niao He +1
In this paper, we study the dynamics of temporal difference learning with neural network-based value function approximation over a general state space, namely, \emph{Neural TD lear…