activity
20172025
most citedTraining data-efficient image transformers & distillation through attention

150 citations · 439 across the 32 of their papers we have counts for

collaborators
Showing cs.LGShow all

10 papers · 1 filter

cs.LG2023

Rewarded soups: towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards

Alexandre Ramé, Guillaume Couairon, Mustafa Shukor +4

Foundation models are first pre-trained on vast unsupervised datasets and then fine-tuned on labeled data. Reinforcement learning, notably from human feedback (RLHF), can further a…

cs.LG2022

Towards efficient feature sharing in MIMO architectures

Rémy Sun, Alexandre Ramé, Clément Masson +2

Multi-input multi-output architectures propose to train multiple subnetworks within one base network and then average the subnetwork predictions to benefit from ensembling for free…

cs.LG2021

RED++ : Data-Free Pruning of Deep Neural Networks via Input Splitting and Output Merging

Edouard Yvinec, Arnaud Dapogny, Matthieu Cord +1

Pruning Deep Neural Networks (DNNs) is a prominent field of study in the goal of inference runtime acceleration. In this paper, we introduce a novel data-free pruning protocol RED+…

cs.LG2021

MixMo: Mixing Multiple Inputs for Multiple Outputs via Deep Subnetworks

Alexandre Rame, Remy Sun, Matthieu Cord

Recent strategies achieved ensembling "for free" by fitting concurrently diverse subnetworks inside a single base network. The main idea during training is that each subnetwork lea…

cs.LG20216 cited

DICE: Diversity in Deep Ensembles via Conditional Redundancy Adversarial Estimation

Alexandre Rame, Matthieu Cord

Deep ensembles perform better than a single network thanks to the diversity among their members. Recent approaches regularize predictions to increase diversity; however, they also…

cs.LG2019

REVE: Regularizing Deep Learning with Variational Entropy Bound

Antoine Saporta, Yifu Chen, Michael Blot +1

Studies on generalization performance of machine learning algorithms under the scope of information theory suggest that compressed representations can guarantee good generalization…