activity
20192025
most citedThe Last-Iterate Convergence Rate of Optimistic Mirror Descent in Stochastic Variational Inequalities

5 citations · 5 across the 1 of their papers we have counts for

collaborators

10 papers

cs.LG2025

How Does the Pretraining Distribution Shape In-Context Learning? A Fundamental Trade-Off

Waïss Azizian, Ali Hasan

The factors driving the performance of in-context learning (ICL) in large language models (LLMs) remain poorly understood despite ICL's surprising effectiveness, enabling models to…

cs.LG2025

The Geometries of Truth Are Orthogonal Across Tasks

Waiss Azizian, Michael Kirchhof, Eugene Ndiaye +4

Large Language Models (LLMs) have demonstrated impressive generalization capabilities across various tasks, but their claim to practical relevance is still mired by concerns on the…

math.OC2025

The global convergence time of stochastic gradient descent in non-convex landscapes: Sharp estimates via large deviations

Waïss Azizian, Franck Iutzeler, Jérôme Malick +1

In this paper, we examine the time it takes for stochastic gradient descent (SGD) to reach the global minimum of a general, non-convex loss function. We approach this question thro…

cs.LG2024

: a library for Wasserstein distributionally robust machine learning

Florian Vincent, Waïss Azizian, Franck Iutzeler +1

We present skwdro, a Python library for training robust machine learning models. The library is based on distributionally robust optimization using Wasserstein distances, popular i…

math.OC2024

What is the long-run distribution of stochastic gradient descent? A large deviations analysis

Waïss Azizian, Franck Iutzeler, Jérôme Malick +1

In this paper, we examine the long-run distribution of stochastic gradient descent (SGD) in general, non-convex problems. Specifically, we seek to understand which regions of the p…

cs.LG2024

Almost sure convergence rates of stochastic gradient methods under gradient domination

Simon Weissmann, Sara Klein, Waïss Azizian +1

Stochastic gradient methods are among the most important algorithms in training machine learning problems. While classical assumptions such as strong convexity allow a simple analy…