activity
20242026
collaborators

7 papers

cs.LG2026

How Does the Pretraining Distribution Shape In-Context Learning? A Fundamental Trade-Off

Waïss Azizian, Ali Hasan

The factors driving the performance of in-context learning (ICL) in large language models (LLMs) remain poorly understood despite ICL's surprising effectiveness, enabling models to…

math.OC2026

What is the long-run distribution of stochastic gradient descent? A large deviations analysis

Waïss Azizian, Franck Iutzeler, Jérôme Malick +1

In this paper, we examine the long-run distribution of stochastic gradient descent (SGD) in general, non-convex problems. Specifically, we seek to understand which regions of the p…

cs.LG2026

: a library for Wasserstein distributionally robust machine learning

Florian Vincent, Waïss Azizian, Franck Iutzeler +1

We present skwdro, a Python library for training robust machine learning models. The library is based on distributionally robust optimization using Wasserstein distances, popular i…

cs.LG2025

The Geometries of Truth Are Orthogonal Across Tasks

Waiss Azizian, Michael Kirchhof, Eugene Ndiaye +4

Large Language Models (LLMs) have demonstrated impressive generalization capabilities across various tasks, but their claim to practical relevance is still mired by concerns on the…

math.OC2025

The global convergence time of stochastic gradient descent in non-convex landscapes: Sharp estimates via large deviations

Waïss Azizian, Franck Iutzeler, Jérôme Malick +1

In this paper, we examine the time it takes for stochastic gradient descent (SGD) to reach the global minimum of a general, non-convex loss function. We approach this question thro…

cs.LG2025

Almost sure convergence rates of stochastic gradient methods under gradient domination

Simon Weissmann, Sara Klein, Waïss Azizian +1

Stochastic gradient methods are among the most important algorithms in training machine learning problems. While classical assumptions such as strong convexity allow a simple analy…