9 citations · 9 across the 1 of their papers we have counts for
3 papers · 1 filter
WARP: On the Benefits of Weight Averaged Rewarded Policies
Alexandre Ramé, Johan Ferret, Nino Vieillard +7
Reinforcement learning from human feedback (RLHF) aligns large language models (LLMs) by encouraging their generations to have high rewards, using a reward model trained on human p…
A Mechanism-Based Approach to Mitigating Harms from Persuasive Generative AI
Seliem El-Sayed, Canfer Akbulut, Amanda McCroskery +17
Recent generative AI systems have demonstrated more advanced persuasive capabilities and are increasingly permeating areas of life where they can influence decision-making. Generat…
DiPaCo: Distributed Path Composition
Arthur Douillard, Qixuan Feng, Andrei A. Rusu +7
Progress in machine learning (ML) has been fueled by scaling neural network models. This scaling has been enabled by ever more heroic feats of engineering, necessary for accommodat…