6 papers
Convergence of projected stochastic natural gradient variational inference for various step size and sample or batch size schedules
Thomas Guilmeau, Hadrien Hendrikx, Florence Forbes
Stochastic natural gradient variational inference (NGVI) is a popular and efficient algorithm for Bayesian inference. Despite empirical success, the convergence of this method is s…
BotaCLIP: Contrastive Learning for Botany-Aware Representation of Earth Observation Data
Selene Cerna, Sara Si-Moussi, Wilfried Thuiller +2
Foundation models have demonstrated a remarkable ability to learn rich, transferable representations across diverse modalities such as images, text, and audio. In modern machine le…
From Inexact Gradients to Byzantine Robustness: Acceleration and Optimization under Similarity
Renaud Gaucher, Aymeric Dieuleveut, Hadrien Hendrikx
Standard federated learning algorithms are vulnerable to adversarial nodes, a.k.a. Byzantine failures. To solve this issue, robust distributed learning algorithms have been develop…
Byzantine-Robust Gossip: Insights from a Dual Approach
Renaud Gaucher, Aymeric Dieuleveut, Hadrien Hendrikx
Distributed learning has many computational benefits but is vulnerable to attacks from a subset of devices transmitting incorrect information. This paper investigates Byzantine-res…
Unified Breakdown Analysis for Byzantine Robust Gossip
Renaud Gaucher, Aymeric Dieuleveut, Hadrien Hendrikx
In decentralized machine learning, different devices communicate in a peer-to-peer manner to collaboratively learn from each other's data. Such approaches are vulnerable to misbeha…
Exponential Moving Average of Weights in Deep Learning: Dynamics and Benefits
Daniel Morales-Brotons, Thijs Vogels, Hadrien Hendrikx
Weight averaging of Stochastic Gradient Descent (SGD) iterates is a popular method for training deep learning models. While it is often used as part of complex training pipelines t…