13 papers
Stabilizing Extrapolation in Looped Transformers via Learned Stochastic Stopping
Hsun-Yu Kuo, El Mahdi Chayti, Patrik Reizinger +2
Looped Transformers, which repeatedly apply a shared transformer block, are an architecturally natural fit for variable-length algorithmic tasks. Although they can exhibit strong l…
Stochastic Zeroth-Order Optimization Under Heavy-Tailed Noise
Taha El Bakkali, El Mahdi Chayti, Qiuyi Zhang +2
We study stochastic zeroth-order (ZO) optimization of smooth nonconvex objectives under heavy-tailed sample-gradient noise. This regime is motivated by empirical evidence that grad…
Stochastic Compositional Optimization via Hybrid Momentum Frank--Wolfe
El Mahdi Chayti
Stochastic compositional optimization minimizes objectives of the form , where is accessible only through noisy st…
RanSOM: Second-Order Momentum with Randomized Scaling for Constrained and Unconstrained Optimization
El Mahdi Chayti
Momentum methods, such as Polyak's Heavy Ball, are the standard for training deep networks but suffer from curvature-induced bias in stochastic settings, limiting convergence to su…
A Split-Client Approach to Second-Order Optimization
El Mahdi Chayti, Martin Jaggi
Second-order optimization methods offer superior convergence rates but are often bottlenecked by the wall-clock cost of Hessian computation and factorization. In the moderate-dimen…
Faster Gradient Methods for Highly-Smooth Stochastic Bilevel Optimization
Lesi Chen, Junru Li, El Mahdi Chayti +1
This paper studies the complexity of finding an -stationary point for stochastic bilevel optimization when the upper-level problem is nonconvex and the lower-level problem is s…