9 citations · 21 across the 5 of their papers we have counts for
4 papers · 1 filter
WARP: On the Benefits of Weight Averaged Rewarded Policies
Alexandre Ramé, Johan Ferret, Nino Vieillard +7
Reinforcement learning from human feedback (RLHF) aligns large language models (LLMs) by encouraging their generations to have high rewards, using a reward model trained on human p…
DiPaCo: Distributed Path Composition
Arthur Douillard, Qixuan Feng, Andrei A. Rusu +7
Progress in machine learning (ML) has been fueled by scaling neural network models. This scaling has been enabled by ever more heroic feats of engineering, necessary for accommodat…
Towards Compute-Optimal Transfer Learning
Massimo Caccia, Alexandre Galashov, Arthur Douillard +6
The field of transfer learning is undergoing a significant shift with the introduction of large pretrained models which have demonstrated strong adaptability to a variety of downst…
Continual Learning with Foundation Models: An Empirical Study of Latent Replay
Oleksiy Ostapenko, Timothee Lesort, Pau Rodríguez +4
Rapid development of large-scale pre-training has resulted in foundation models that can act as effective feature extractors on a variety of downstream tasks and domains. Motivated…