papers

Publications (44)

cs.LG2020

Amazon SageMaker Autopilot: a white box AutoML solution at scale

Piali Das, Valerio Perrone, Nikita Ivkin +22

AutoML systems provide a black-box solution to machine learning problems by selecting the right way of processing features, choosing an algorithm and tuning the hyperparameters of…

stat.ML2015

Sample Complexity of Dictionary Learning and other Matrix Factorizations

Rémi Gribonval, Rodolphe Jenatton, Francis Bach +2

Many modern tools in machine learning and signal processing, such as sparse dictionary learning, principal component analysis (PCA), non-negative matrix factorization (NMF), -me…

cs.LG2021

Predicting the utility of search spaces for black-box optimization: a simple, budget-aware approach

Setareh Ariafar, Justin Gilmer, Zachary Nado +3

Black box optimization requires specifying a search space to explore for solutions, e.g. a d-dimensional compact space, and this choice is critical for getting the best results at…

cs.LG2011

Learning Hierarchical and Topographic Dictionaries with Structured Sparsity

Julien Mairal, Rodolphe Jenatton, Guillaume Obozinski +1

Recent work in signal processing and statistics have focused on defining new regularization functions, which not only induce sparsity of the solution, but also take into account th…

cs.CV2021

Scaling Vision with Sparse Mixture of Experts

Carlos Riquelme, Joan Puigcerver, Basil Mustafa +5

Sparsely-gated Mixture of Experts networks (MoEs) have demonstrated excellent scalability in Natural Language Processing. In Computer Vision, however, almost all performant network…

cs.CV2022

Multimodal Contrastive Learning with LIMoE: the Language-Image Mixture of Experts

Basil Mustafa, Carlos Riquelme, Joan Puigcerver +2

Large sparsely-activated models have obtained excellent performance in multiple domains. However, such models are typically trained on a single modality at a time. We present the L…

stat.ML2022

Deep Classifiers with Label Noise Modeling and Distance Awareness

Vincent Fortuin, Mark Collier, Florian Wenzel +7

Uncertainty estimation in deep learning has recently emerged as a crucial area of interest to advance reliability and robustness in safety-critical applications. While there have b…

cs.LG2021

Amazon SageMaker Automatic Model Tuning: Scalable Gradient-Free Optimization

Valerio Perrone, Huibin Shen, Aida Zolic +12

Tuning complex machine learning systems is challenging. Machine learning typically requires to set hyperparameters, be it regularization, architecture, or optimization parameters,…

stat.ML2016

Online optimization and regret guarantees for non-additive long-term constraints

Rodolphe Jenatton, Jim Huang, Dominik Csiba +1

We consider online optimization in the 1-lookahead setting, where the objective does not decompose additively over the rounds of the online game. The resulting formulation enables…

stat.ML2015

Adaptive Algorithms for Online Convex Optimization with Long-term Constraints

Rodolphe Jenatton, Jim Huang, Cédric Archambeau

We present an adaptive online gradient descent algorithm to solve online convex optimization problems with long-term constraints , which are constraints that need to be satisfied w…

cs.LG2022

Uncertainty Baselines: Benchmarks for Uncertainty & Robustness in Deep Learning

Zachary Nado, Neil Band, Mark Collier +23

High-quality estimates of uncertainty and robustness are crucial for numerous real-world applications, especially for deep learning which underlies many deployed ML systems. The ab…

cs.LG2022

On Mixup Regularization

Luigi Carratino, Moustapha Cissé, Rodolphe Jenatton +1

Mixup is a data augmentation technique that creates new examples as convex combinations of training points and labels. This simple technique has empirically shown to improve the ac…

cs.LG2022

On the Adversarial Robustness of Mixture of Experts

Joan Puigcerver, Rodolphe Jenatton, Carlos Riquelme +2

Adversarial robustness is a key desirable property of neural networks. It has been empirically shown to be affected by their sizes, with larger networks being typically more robust…

cs.LG2010

Network Flow Algorithms for Structured Sparsity

Julien Mairal, Rodolphe Jenatton, Guillaume Obozinski +1

We consider a class of learning problems that involve a structured sparsity-inducing norm defined as the sum of -norms over groups of variables. Whereas a lot of effor…

stat.ML2012

Local stability and robustness of sparse dictionary learning in the presence of noise

Rodolphe Jenatton, Rémi Gribonval, Francis Bach

A popular approach within the signal processing and machine learning communities consists in modelling signals as sparse linear combinations of atoms selected from a learned dictio…

cs.CV2023

Three Towers: Flexible Contrastive Learning with Pretrained Image Models

Jannik Kossen, Mark Collier, Basil Mustafa +7

We introduce Three Towers (3T), a flexible method to improve the contrastive learning of vision-language models by incorporating pretrained image classifiers. While contrastive mod…

cs.LG2021

Hydra: Preserving Ensemble Diversity for Model Distillation

Linh Tran, Bastiaan S. Veeling, Kevin Roth +7

Ensembles of models have been empirically shown to improve predictive performance and to yield robust measures of uncertainty. However, they are expensive in computation and memory…

math.OC2011

Convex and Network Flow Optimization for Structured Sparsity

Julien Mairal, Rodolphe Jenatton, Guillaume Obozinski +1

We consider a class of learning problems regularized by a structured sparsity-inducing norm defined as the sum of l_2- or l_infinity-norms over groups of variables. Whereas much ef…

cs.LG2015

Sparse and spurious: dictionary learning with noise and outliers

Rémi Gribonval, Rodolphe Jenatton, Francis Bach

A popular approach within the signal processing and machine learning communities consists in modelling signals as sparse linear combinations of atoms selected from a learned dictio…

cs.LG2022

Transfer and Marginalize: Explaining Away Label Noise with Privileged Information

Mark Collier, Rodolphe Jenatton, Efi Kokiopoulou +1

Supervised learning datasets often have privileged information, in the form of features which are available at training time but are not available at test time e.g. the ID of the a…

stat.ML2014

On The Sample Complexity of Sparse Dictionary Learning

Matthias Seibert, Martin Kleinsteuber, Rémi Gribonval +2

In the synthesis model signals are represented as a sparse combinations of atoms from a dictionary. Dictionary learning describes the acquisition process of the underlying dictiona…

cs.LG2011

Optimization with Sparsity-Inducing Penalties

Francis Bach, Rodolphe Jenatton, Julien Mairal +1

Sparse estimation methods are aimed at using or obtaining parsimonious representations of data or models. They were first dedicated to linear variable selection but numerous extens…

cs.LG2023

Massively Scaling Heteroscedastic Classifiers

Mark Collier, Rodolphe Jenatton, Basil Mustafa +3

Heteroscedastic classifiers, which learn a multivariate Gaussian distribution over prediction logits, have been shown to perform well on image classification problems with hundreds…

stat.ML2019

Constrained Bayesian Optimization with Max-Value Entropy Search

Valerio Perrone, Iaroslav Shcherbatyi, Rodolphe Jenatton +2

Bayesian optimization (BO) is a model-based approach to sequentially optimize expensive black-box functions, such as the validation error of a deep neural network with respect to i…

stat.ML2011

Multi-scale Mining of fMRI data with Hierarchical Structured Sparsity

Rodolphe Jenatton, Alexandre Gramfort, Vincent Michel +4

Inverse inference, or "brain reading", is a recent paradigm for analyzing functional magnetic resonance imaging (fMRI) data, based on pattern recognition and statistical learning.…

cs.LG2021

Training independent subnetworks for robust prediction

Marton Havasi, Rodolphe Jenatton, Stanislav Fort +5

Recent approaches to efficiently ensemble neural networks have shown that strong robustness and uncertainty performance can be achieved with a negligible gain in parameters over th…

cs.LG2022

Plex: Towards Reliability using Pretrained Large Model Extensions

Dustin Tran, Jeremiah Liu, Michael W. Dusenberry +23

A recent trend in artificial intelligence is the use of pretrained models for language and vision tasks, which have achieved extraordinary performance but also puzzling failures. P…

cs.CL2025

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Gheorghe Comanici, Eric Bieber, Mike Schaekermann +3431

In this report, we introduce the Gemini 2.X model family: Gemini 2.5 Pro and Gemini 2.5 Flash, as well as our earlier Gemini 2.0 Flash and Flash-Lite models. Gemini 2.5 Pro is our…

math.OC2015

Convex Relaxations for Permutation Problems

Fajwel Fogel, Rodolphe Jenatton, Francis Bach +1

Seriation seeks to reconstruct a linear order between variables using unsorted, pairwise similarity information. It has direct applications in archeology and shotgun gene sequencin…

stat.ML2019

Learning search spaces for Bayesian optimization: Another view of hyperparameter transfer learning

Valerio Perrone, Huibin Shen, Matthias Seeger +2

Bayesian optimization (BO) is a successful methodology to optimize black-box functions that are expensive to evaluate. While traditional methods optimize each black-box function in…

cs.LG2024

Pi-DUAL: Using Privileged Information to Distinguish Clean from Noisy Labels

Ke Wang, Guillermo Ortiz-Jimenez, Rodolphe Jenatton +3

Label noise is a pervasive problem in deep learning that often compromises the generalization performance of trained models. Recently, leveraging privileged information (PI) -- inf…

stat.ML2011

Proximal Methods for Hierarchical Sparse Coding

Rodolphe Jenatton, Julien Mairal, Guillaume Obozinski +1

Sparse coding consists in representing signals as sparse linear combinations of atoms selected from a dictionary. We consider an extension of this framework where the atoms are fur…

stat.ML2020

How Good is the Bayes Posterior in Deep Neural Networks Really?

Florian Wenzel, Kevin Roth, Bastiaan S. Veeling +7

During the past five years the Bayesian deep learning community has developed increasingly accurate and efficient approximate inference procedures that allow for Bayesian inference…

cs.CV2023

Scaling Vision Transformers to 22 Billion Parameters

Mostafa Dehghani, Josip Djolonga, Basil Mustafa +39

The scaling of Transformers has driven breakthrough capabilities for language models. At present, the largest large language models (LLMs) contain upwards of 100B parameters. Visio…

cs.LG2020

The k-tied Normal Distribution: A Compact Parameterization of Gaussian Mean Field Posteriors in Bayesian Neural Networks

Jakub Swiatkowski, Kevin Roth, Bastiaan S. Veeling +7

Variational Bayesian Inference is a popular methodology for approximating posterior distributions over Bayesian neural network weights. Recent work developing this class of methods…

cs.LG2021

Correlated Input-Dependent Label Noise in Large-Scale Image Classification

Mark Collier, Basil Mustafa, Efi Kokiopoulou +2

Large scale image classification datasets often contain noisy labels. We take a principled probabilistic approach to modelling input-dependent, also known as heteroscedastic, label…

stat.ML2017

Multiple Adaptive Bayesian Linear Regression for Scalable Bayesian Optimization with Warm Start

Valerio Perrone, Rodolphe Jenatton, Matthias Seeger +1

Bayesian optimization (BO) is a model-based approach for gradient-free black-box function optimization. Typically, BO is powered by a Gaussian process (GP), whose algorithmic compl…

cs.LG2023

When does Privileged Information Explain Away Label Noise?

Guillermo Ortiz-Jimenez, Mark Collier, Anant Nawalgaria +4

Leveraging privileged information (PI), or features available during training but not at test time, has recently been shown to be an effective method for addressing label noise. Ho…

cs.LG2023

Sparse MoEs meet Efficient Ensembles

James Urquhart Allingham, Florian Wenzel, Zelda E Mariet +10

Machine learning models based on the aggregated outputs of submodels, either at the activation or prediction levels, often exhibit strong performance compared to individual models.…

cs.LG2012

Structured sparsity through convex optimization

Francis Bach, Rodolphe Jenatton, Julien Mairal +1

Sparse estimation methods are aimed at using or obtaining parsimonious representations of data or models. While naturally cast as a combinatorial optimization problem, variable or…

stat.ML2010

Structured Variable Selection with Sparsity-Inducing Norms

Rodolphe Jenatton, Jean-Yves Audibert, Francis Bach

We consider the empirical risk minimization problem for linear supervised learning, with regularization by structured sparsity-inducing norms. These are defined as sums of Euclidea…

cs.LG2020

A Simple Probabilistic Method for Deep Classification under Input-Dependent Label Noise

Mark Collier, Basil Mustafa, Efi Kokiopoulou +2

Datasets with noisy labels are a common occurrence in practical applications of classification methods. We propose a simple probabilistic method for training deep classifiers under…

cs.LG2021

Hyperparameter Ensembles for Robustness and Uncertainty Quantification

Florian Wenzel, Jasper Snoek, Dustin Tran +1

Ensembles over neural network weights trained from different random initialization, known as deep ensembles, achieve state-of-the-art accuracy and calibration. The recently introdu…

stat.ML2009

Structured Sparse Principal Component Analysis

Rodolphe Jenatton, Guillaume Obozinski, Francis Bach

We present an extension of sparse PCA, or sparse dictionary learning, where the sparsity patterns of all dictionary elements are structured and constrained to belong to a prespecif…