NewEvery arXiv paper, its researchers & institutions — mapped.
papers

Publications (27)

cs.LG2024

MusicRL: Aligning Music Generation to Human Preferences

Geoffrey Cideron, Sertan Girgin, Mauro Verzetti +11

cs.LG2021

What Matters for Adversarial Imitation Learning?

Manu Orsini, Anton Raichuk, Léonard Hussenot +7

cs.LG2024

RecurrentGemma: Moving Past Transformers for Efficient Open Language Models

Aleksandar Botev, Soham De, Samuel L Smith +59

cs.CL2024

Gemma 2: Improving Open Language Models at a Practical Size

Gemma Team, Morgane Riviere, Shreya Pathak +195

math.DS2021

Solving N-player dynamic routing games with congestion: a mean field approach

Theophile Cabannes, Mathieu Lauriere, Julien Perolat +7

stat.ML2024

Nash Learning from Human Feedback

Rémi Munos, Michal Valko, Daniele Calandriello +14

cs.CL2023

Factually Consistent Summarization via Reinforcement Learning with Textual Entailment Feedback

Paul Roit, Johan Ferret, Lior Shani +16

cs.LG2020

What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study

Marcin Andrychowicz, Anton Raichuk, Piotr Stańczyk +9

cs.LG2024

BOND: Aligning LLMs with Best-of-N Distillation

Pier Giuseppe Sessa, Robert Dadashi, Léonard Hussenot +17

cs.CL2022

Decoding a Neural Retriever's Latent Space for Query Suggestion

Leonard Adolphs, Michelle Chen Huebscher, Christian Buck +4

cs.CL2026

DiffusionGemma Technical Report

DiffusionGemma Team, Adrien Ali Taïga, James Assiene +41

cs.RO2021

Brax -- A Differentiable Physics Engine for Large Scale Rigid Body Simulation

C. Daniel Freeman, Erik Frey, Anton Raichuk +3

cs.CL2025

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Gheorghe Comanici, Eric Bieber, Mike Schaekermann +3431

cs.LG2024

WARP: On the Benefits of Weight Averaged Rewarded Policies

Alexandre Ramé, Johan Ferret, Nino Vieillard +7

cs.SD2023

Speak, Read and Prompt: High-Fidelity Text-to-Speech with Minimal Supervision

Eugene Kharitonov, Damien Vincent, Zalán Borsos +6

cs.LG2024

Learning in Mean Field Games: A Survey

Mathieu Laurière, Sarah Perrin, Julien Pérolat +5

cs.LG2022

Acme: A Research Framework for Distributed Reinforcement Learning

Matthew W. Hoffman, Bobak Shahriari, John Aslanides +36

cs.CL2022

vec2text with Round-Trip Translations

Geoffrey Cideron, Sertan Girgin, Anton Raichuk +3

cs.RO2023

Get Back Here: Robust Imitation by Return-to-Distribution Planning

Geoffrey Cideron, Baruch Tabanpour, Sebastian Curi +6

cs.LG2021

RLDS: an Ecosystem to Generate, Share and Use Datasets in Reinforcement Learning

Sabela Ramos, Sertan Girgin, Léonard Hussenot +9

cs.LG2022

Scalable Deep Reinforcement Learning Algorithms for Mean Field Games

Mathieu Laurière, Sarah Perrin, Sertan Girgin +8

cs.CL2025

Gemma 3 Technical Report

Gemma Team, Aishwarya Kamath, Johan Ferret +209

cs.CL2026

Gemma 4 Technical Report

Gemma Team, Sherif El Abd, Vaibhav Aggarwal +320

cs.LG2024

Diversity-Rewarded CFG Distillation

Geoffrey Cideron, Andrea Agostinelli, Johan Ferret +5

cs.CL2024

Gemma: Open Models Based on Gemini Research and Technology

Gemma Team, Thomas Mesnard, Cassidy Hardin +105

cs.LG2022

Continuous Control with Action Quantization from Demonstrations

Robert Dadashi, Léonard Hussenot, Damien Vincent +4

cs.LG2021

Hyperparameter Selection for Imitation Learning

Leonard Hussenot, Marcin Andrychowicz, Damien Vincent +11