activity
20182026
most citedGroup-Fair Online Allocation in Continuous Time

11 citations · 20 across the 9 of their papers we have counts for

collaborators
Showing cs.LGShow all

9 papers · 1 filter

cs.LG2026

Finite-Time Analysis of Gradient Descent for Shallow Transformers

Enes Arda, Semih Cayci, Atilla Eryilmaz

Understanding why Transformers perform so well remains challenging due to their non-convex optimization landscape. In this work, we analyze a shallow Transformer with independe…

cs.LG2025

Non-Asymptotic Optimization and Generalization Bounds for Stochastic Gauss-Newton in Overparameterized Models

Semih Cayci

An important question in deep learning is how higher-order optimization methods affect generalization. In this work, we analyze a stochastic Gauss-Newton (SGN) method with Levenber…

cs.LG2025

Convergence of Stochastic Gradient Langevin Dynamics in the Lazy Training Regime

Noah Oberweis, Semih Cayci

Continuous-time models provide important insights into the training dynamics of optimization algorithms in deep learning. In this work, we establish a non-asymptotic convergence an…

cs.LG2024

Convergence of Gradient Descent for Recurrent Neural Networks: A Nonasymptotic Analysis

Semih Cayci, Atilla Eryilmaz

We analyze recurrent neural networks with diagonal hidden-to-hidden weight matrices, trained with gradient descent in the supervised learning setting, and prove that gradient desce…

cs.LG2023

Provably Robust Temporal Difference Learning for Heavy-Tailed Rewards

Semih Cayci, Atilla Eryilmaz

In a broad class of reinforcement learning applications, stochastic rewards have heavy-tailed distributions, which lead to infinite second-order moments for stochastic (semi)gradie…

cs.LG2021

Sample Complexity and Overparameterization Bounds for Temporal Difference Learning with Neural Network Approximation

Semih Cayci, Siddhartha Satpathi, Niao He +1

In this paper, we study the dynamics of temporal difference learning with neural network-based value function approximation over a general state space, namely, \emph{Neural TD lear…