Publications (44)
Amazon SageMaker Autopilot: a white box AutoML solution at scale
Piali Das, Valerio Perrone, Nikita Ivkin +22
AutoML systems provide a black-box solution to machine learning problems by selecting the right way of processing features, choosing an algorithm and tuning the hyperparameters of…
Sample Complexity of Dictionary Learning and other Matrix Factorizations
Rémi Gribonval, Rodolphe Jenatton, Francis Bach +2
Many modern tools in machine learning and signal processing, such as sparse dictionary learning, principal component analysis (PCA), non-negative matrix factorization (NMF), -me…
Predicting the utility of search spaces for black-box optimization: a simple, budget-aware approach
Setareh Ariafar, Justin Gilmer, Zachary Nado +3
Black box optimization requires specifying a search space to explore for solutions, e.g. a d-dimensional compact space, and this choice is critical for getting the best results at…
Learning Hierarchical and Topographic Dictionaries with Structured Sparsity
Julien Mairal, Rodolphe Jenatton, Guillaume Obozinski +1
Recent work in signal processing and statistics have focused on defining new regularization functions, which not only induce sparsity of the solution, but also take into account th…
Scaling Vision with Sparse Mixture of Experts
Carlos Riquelme, Joan Puigcerver, Basil Mustafa +5
Sparsely-gated Mixture of Experts networks (MoEs) have demonstrated excellent scalability in Natural Language Processing. In Computer Vision, however, almost all performant network…
Multimodal Contrastive Learning with LIMoE: the Language-Image Mixture of Experts
Basil Mustafa, Carlos Riquelme, Joan Puigcerver +2
Large sparsely-activated models have obtained excellent performance in multiple domains. However, such models are typically trained on a single modality at a time. We present the L…
Deep Classifiers with Label Noise Modeling and Distance Awareness
Vincent Fortuin, Mark Collier, Florian Wenzel +7
Uncertainty estimation in deep learning has recently emerged as a crucial area of interest to advance reliability and robustness in safety-critical applications. While there have b…
Amazon SageMaker Automatic Model Tuning: Scalable Gradient-Free Optimization
Valerio Perrone, Huibin Shen, Aida Zolic +12
Tuning complex machine learning systems is challenging. Machine learning typically requires to set hyperparameters, be it regularization, architecture, or optimization parameters,…
Online optimization and regret guarantees for non-additive long-term constraints
Rodolphe Jenatton, Jim Huang, Dominik Csiba +1
We consider online optimization in the 1-lookahead setting, where the objective does not decompose additively over the rounds of the online game. The resulting formulation enables…
Adaptive Algorithms for Online Convex Optimization with Long-term Constraints
Rodolphe Jenatton, Jim Huang, Cédric Archambeau
We present an adaptive online gradient descent algorithm to solve online convex optimization problems with long-term constraints , which are constraints that need to be satisfied w…
Uncertainty Baselines: Benchmarks for Uncertainty & Robustness in Deep Learning
Zachary Nado, Neil Band, Mark Collier +23
High-quality estimates of uncertainty and robustness are crucial for numerous real-world applications, especially for deep learning which underlies many deployed ML systems. The ab…
On Mixup Regularization
Luigi Carratino, Moustapha Cissé, Rodolphe Jenatton +1
Mixup is a data augmentation technique that creates new examples as convex combinations of training points and labels. This simple technique has empirically shown to improve the ac…
On the Adversarial Robustness of Mixture of Experts
Joan Puigcerver, Rodolphe Jenatton, Carlos Riquelme +2
Adversarial robustness is a key desirable property of neural networks. It has been empirically shown to be affected by their sizes, with larger networks being typically more robust…
Network Flow Algorithms for Structured Sparsity
Julien Mairal, Rodolphe Jenatton, Guillaume Obozinski +1
We consider a class of learning problems that involve a structured sparsity-inducing norm defined as the sum of -norms over groups of variables. Whereas a lot of effor…
Local stability and robustness of sparse dictionary learning in the presence of noise
Rodolphe Jenatton, Rémi Gribonval, Francis Bach
A popular approach within the signal processing and machine learning communities consists in modelling signals as sparse linear combinations of atoms selected from a learned dictio…
Three Towers: Flexible Contrastive Learning with Pretrained Image Models
Jannik Kossen, Mark Collier, Basil Mustafa +7
We introduce Three Towers (3T), a flexible method to improve the contrastive learning of vision-language models by incorporating pretrained image classifiers. While contrastive mod…
Hydra: Preserving Ensemble Diversity for Model Distillation
Linh Tran, Bastiaan S. Veeling, Kevin Roth +7
Ensembles of models have been empirically shown to improve predictive performance and to yield robust measures of uncertainty. However, they are expensive in computation and memory…
Convex and Network Flow Optimization for Structured Sparsity
Julien Mairal, Rodolphe Jenatton, Guillaume Obozinski +1
We consider a class of learning problems regularized by a structured sparsity-inducing norm defined as the sum of l_2- or l_infinity-norms over groups of variables. Whereas much ef…
Sparse and spurious: dictionary learning with noise and outliers
Rémi Gribonval, Rodolphe Jenatton, Francis Bach
A popular approach within the signal processing and machine learning communities consists in modelling signals as sparse linear combinations of atoms selected from a learned dictio…
Transfer and Marginalize: Explaining Away Label Noise with Privileged Information
Mark Collier, Rodolphe Jenatton, Efi Kokiopoulou +1
Supervised learning datasets often have privileged information, in the form of features which are available at training time but are not available at test time e.g. the ID of the a…
On The Sample Complexity of Sparse Dictionary Learning
Matthias Seibert, Martin Kleinsteuber, Rémi Gribonval +2
In the synthesis model signals are represented as a sparse combinations of atoms from a dictionary. Dictionary learning describes the acquisition process of the underlying dictiona…
Optimization with Sparsity-Inducing Penalties
Francis Bach, Rodolphe Jenatton, Julien Mairal +1
Sparse estimation methods are aimed at using or obtaining parsimonious representations of data or models. They were first dedicated to linear variable selection but numerous extens…
Massively Scaling Heteroscedastic Classifiers
Mark Collier, Rodolphe Jenatton, Basil Mustafa +3
Heteroscedastic classifiers, which learn a multivariate Gaussian distribution over prediction logits, have been shown to perform well on image classification problems with hundreds…
Constrained Bayesian Optimization with Max-Value Entropy Search
Valerio Perrone, Iaroslav Shcherbatyi, Rodolphe Jenatton +2
Bayesian optimization (BO) is a model-based approach to sequentially optimize expensive black-box functions, such as the validation error of a deep neural network with respect to i…
Multi-scale Mining of fMRI data with Hierarchical Structured Sparsity
Rodolphe Jenatton, Alexandre Gramfort, Vincent Michel +4
Inverse inference, or "brain reading", is a recent paradigm for analyzing functional magnetic resonance imaging (fMRI) data, based on pattern recognition and statistical learning.…
Training independent subnetworks for robust prediction
Marton Havasi, Rodolphe Jenatton, Stanislav Fort +5
Recent approaches to efficiently ensemble neural networks have shown that strong robustness and uncertainty performance can be achieved with a negligible gain in parameters over th…
Plex: Towards Reliability using Pretrained Large Model Extensions
Dustin Tran, Jeremiah Liu, Michael W. Dusenberry +23
A recent trend in artificial intelligence is the use of pretrained models for language and vision tasks, which have achieved extraordinary performance but also puzzling failures. P…
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Gheorghe Comanici, Eric Bieber, Mike Schaekermann +3431
In this report, we introduce the Gemini 2.X model family: Gemini 2.5 Pro and Gemini 2.5 Flash, as well as our earlier Gemini 2.0 Flash and Flash-Lite models. Gemini 2.5 Pro is our…
Convex Relaxations for Permutation Problems
Fajwel Fogel, Rodolphe Jenatton, Francis Bach +1
Seriation seeks to reconstruct a linear order between variables using unsorted, pairwise similarity information. It has direct applications in archeology and shotgun gene sequencin…
Learning search spaces for Bayesian optimization: Another view of hyperparameter transfer learning
Valerio Perrone, Huibin Shen, Matthias Seeger +2
Bayesian optimization (BO) is a successful methodology to optimize black-box functions that are expensive to evaluate. While traditional methods optimize each black-box function in…
Pi-DUAL: Using Privileged Information to Distinguish Clean from Noisy Labels
Ke Wang, Guillermo Ortiz-Jimenez, Rodolphe Jenatton +3
Label noise is a pervasive problem in deep learning that often compromises the generalization performance of trained models. Recently, leveraging privileged information (PI) -- inf…
Proximal Methods for Hierarchical Sparse Coding
Rodolphe Jenatton, Julien Mairal, Guillaume Obozinski +1
Sparse coding consists in representing signals as sparse linear combinations of atoms selected from a dictionary. We consider an extension of this framework where the atoms are fur…
How Good is the Bayes Posterior in Deep Neural Networks Really?
Florian Wenzel, Kevin Roth, Bastiaan S. Veeling +7
During the past five years the Bayesian deep learning community has developed increasingly accurate and efficient approximate inference procedures that allow for Bayesian inference…
Scaling Vision Transformers to 22 Billion Parameters
Mostafa Dehghani, Josip Djolonga, Basil Mustafa +39
The scaling of Transformers has driven breakthrough capabilities for language models. At present, the largest large language models (LLMs) contain upwards of 100B parameters. Visio…
The k-tied Normal Distribution: A Compact Parameterization of Gaussian Mean Field Posteriors in Bayesian Neural Networks
Jakub Swiatkowski, Kevin Roth, Bastiaan S. Veeling +7
Variational Bayesian Inference is a popular methodology for approximating posterior distributions over Bayesian neural network weights. Recent work developing this class of methods…
Correlated Input-Dependent Label Noise in Large-Scale Image Classification
Mark Collier, Basil Mustafa, Efi Kokiopoulou +2
Large scale image classification datasets often contain noisy labels. We take a principled probabilistic approach to modelling input-dependent, also known as heteroscedastic, label…
Multiple Adaptive Bayesian Linear Regression for Scalable Bayesian Optimization with Warm Start
Valerio Perrone, Rodolphe Jenatton, Matthias Seeger +1
Bayesian optimization (BO) is a model-based approach for gradient-free black-box function optimization. Typically, BO is powered by a Gaussian process (GP), whose algorithmic compl…
When does Privileged Information Explain Away Label Noise?
Guillermo Ortiz-Jimenez, Mark Collier, Anant Nawalgaria +4
Leveraging privileged information (PI), or features available during training but not at test time, has recently been shown to be an effective method for addressing label noise. Ho…
Sparse MoEs meet Efficient Ensembles
James Urquhart Allingham, Florian Wenzel, Zelda E Mariet +10
Machine learning models based on the aggregated outputs of submodels, either at the activation or prediction levels, often exhibit strong performance compared to individual models.…
Structured sparsity through convex optimization
Francis Bach, Rodolphe Jenatton, Julien Mairal +1
Sparse estimation methods are aimed at using or obtaining parsimonious representations of data or models. While naturally cast as a combinatorial optimization problem, variable or…
Structured Variable Selection with Sparsity-Inducing Norms
Rodolphe Jenatton, Jean-Yves Audibert, Francis Bach
We consider the empirical risk minimization problem for linear supervised learning, with regularization by structured sparsity-inducing norms. These are defined as sums of Euclidea…
A Simple Probabilistic Method for Deep Classification under Input-Dependent Label Noise
Mark Collier, Basil Mustafa, Efi Kokiopoulou +2
Datasets with noisy labels are a common occurrence in practical applications of classification methods. We propose a simple probabilistic method for training deep classifiers under…
Hyperparameter Ensembles for Robustness and Uncertainty Quantification
Florian Wenzel, Jasper Snoek, Dustin Tran +1
Ensembles over neural network weights trained from different random initialization, known as deep ensembles, achieve state-of-the-art accuracy and calibration. The recently introdu…
Structured Sparse Principal Component Analysis
Rodolphe Jenatton, Guillaume Obozinski, Francis Bach
We present an extension of sparse PCA, or sparse dictionary learning, where the sparsity patterns of all dictionary elements are structured and constrained to belong to a prespecif…