7 papers
How Does the Pretraining Distribution Shape In-Context Learning? A Fundamental Trade-Off
Waïss Azizian, Ali Hasan
The factors driving the performance of in-context learning (ICL) in large language models (LLMs) remain poorly understood despite ICL's surprising effectiveness, enabling models to…
What is the long-run distribution of stochastic gradient descent? A large deviations analysis
Waïss Azizian, Franck Iutzeler, Jérôme Malick +1
In this paper, we examine the long-run distribution of stochastic gradient descent (SGD) in general, non-convex problems. Specifically, we seek to understand which regions of the p…
: a library for Wasserstein distributionally robust machine learning
Florian Vincent, Waïss Azizian, Franck Iutzeler +1
We present skwdro, a Python library for training robust machine learning models. The library is based on distributionally robust optimization using Wasserstein distances, popular i…
The Geometries of Truth Are Orthogonal Across Tasks
Waiss Azizian, Michael Kirchhof, Eugene Ndiaye +4
Large Language Models (LLMs) have demonstrated impressive generalization capabilities across various tasks, but their claim to practical relevance is still mired by concerns on the…
The global convergence time of stochastic gradient descent in non-convex landscapes: Sharp estimates via large deviations
Waïss Azizian, Franck Iutzeler, Jérôme Malick +1
In this paper, we examine the time it takes for stochastic gradient descent (SGD) to reach the global minimum of a general, non-convex loss function. We approach this question thro…
Almost sure convergence rates of stochastic gradient methods under gradient domination
Simon Weissmann, Sara Klein, Waïss Azizian +1
Stochastic gradient methods are among the most important algorithms in training machine learning problems. While classical assumptions such as strong convexity allow a simple analy…