Publications (27)
MusicRL: Aligning Music Generation to Human Preferences
Geoffrey Cideron, Sertan Girgin, Mauro Verzetti +11
What Matters for Adversarial Imitation Learning?
Manu Orsini, Anton Raichuk, Léonard Hussenot +7
RecurrentGemma: Moving Past Transformers for Efficient Open Language Models
Aleksandar Botev, Soham De, Samuel L Smith +59
Gemma 2: Improving Open Language Models at a Practical Size
Gemma Team, Morgane Riviere, Shreya Pathak +195
Solving N-player dynamic routing games with congestion: a mean field approach
Theophile Cabannes, Mathieu Lauriere, Julien Perolat +7
Nash Learning from Human Feedback
Rémi Munos, Michal Valko, Daniele Calandriello +14
Factually Consistent Summarization via Reinforcement Learning with Textual Entailment Feedback
Paul Roit, Johan Ferret, Lior Shani +16
What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study
Marcin Andrychowicz, Anton Raichuk, Piotr StaÅczyk +9
BOND: Aligning LLMs with Best-of-N Distillation
Pier Giuseppe Sessa, Robert Dadashi, Léonard Hussenot +17
Decoding a Neural Retriever's Latent Space for Query Suggestion
Leonard Adolphs, Michelle Chen Huebscher, Christian Buck +4
DiffusionGemma Technical Report
DiffusionGemma Team, Adrien Ali Taïga, James Assiene +41
Brax -- A Differentiable Physics Engine for Large Scale Rigid Body Simulation
C. Daniel Freeman, Erik Frey, Anton Raichuk +3
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Gheorghe Comanici, Eric Bieber, Mike Schaekermann +3431
WARP: On the Benefits of Weight Averaged Rewarded Policies
Alexandre Ramé, Johan Ferret, Nino Vieillard +7
Speak, Read and Prompt: High-Fidelity Text-to-Speech with Minimal Supervision
Eugene Kharitonov, Damien Vincent, Zalán Borsos +6
Learning in Mean Field Games: A Survey
Mathieu Laurière, Sarah Perrin, Julien Pérolat +5
Acme: A Research Framework for Distributed Reinforcement Learning
Matthew W. Hoffman, Bobak Shahriari, John Aslanides +36
vec2text with Round-Trip Translations
Geoffrey Cideron, Sertan Girgin, Anton Raichuk +3
Get Back Here: Robust Imitation by Return-to-Distribution Planning
Geoffrey Cideron, Baruch Tabanpour, Sebastian Curi +6
RLDS: an Ecosystem to Generate, Share and Use Datasets in Reinforcement Learning
Sabela Ramos, Sertan Girgin, Léonard Hussenot +9
Scalable Deep Reinforcement Learning Algorithms for Mean Field Games
Mathieu Laurière, Sarah Perrin, Sertan Girgin +8
Gemma 3 Technical Report
Gemma Team, Aishwarya Kamath, Johan Ferret +209
Gemma 4 Technical Report
Gemma Team, Sherif El Abd, Vaibhav Aggarwal +320
Diversity-Rewarded CFG Distillation
Geoffrey Cideron, Andrea Agostinelli, Johan Ferret +5
Gemma: Open Models Based on Gemini Research and Technology
Gemma Team, Thomas Mesnard, Cassidy Hardin +105
Continuous Control with Action Quantization from Demonstrations
Robert Dadashi, Léonard Hussenot, Damien Vincent +4
Hyperparameter Selection for Imitation Learning
Leonard Hussenot, Marcin Andrychowicz, Damien Vincent +11