NewEvery arXiv paper, its researchers & institutions — mapped.
papers

Publications (26)

cs.LG2022

Learning Energy Networks with Generalized Fenchel-Young Losses

Mathieu Blondel, Felipe Llinares-López, Robert Dadashi +2

cs.LG2021

Offline Reinforcement Learning as Anti-Exploration

Shideh Rezaeifar, Robert Dadashi, Nino Vieillard +4

cs.LG2021

Offline Reinforcement Learning with Pseudometric Learning

Robert Dadashi, Shideh Rezaeifar, Nino Vieillard +3

cs.CL2025

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Gheorghe Comanici, Eric Bieber, Mike Schaekermann +3431

cs.CL2025

Gemini: A Family of Highly Capable Multimodal Models

Gemini Team, Rohan Anil, Sebastian Borgeaud +1340

cs.LG2024

WARP: On the Benefits of Weight Averaged Rewarded Policies

Alexandre Ramé, Johan Ferret, Nino Vieillard +7

cs.LG2022

Acme: A Research Framework for Distributed Reinforcement Learning

Matthew W. Hoffman, Bobak Shahriari, John Aslanides +36

cs.CL2022

vec2text with Round-Trip Translations

Geoffrey Cideron, Sertan Girgin, Anton Raichuk +3

cs.LG2021

RLDS: an Ecosystem to Generate, Share and Use Datasets in Reinforcement Learning

Sabela Ramos, Sertan Girgin, Léonard Hussenot +9

cs.CL2025

Gemma 3 Technical Report

Gemma Team, Aishwarya Kamath, Johan Ferret +209

cs.LG2021

Show me the Way: Intrinsic Motivation from Demonstrations

Léonard Hussenot, Robert Dadashi, Matthieu Geist +1

cs.LG2024

WARM: On the Benefits of Weight Averaged Reward Models

Alexandre Ramé, Nino Vieillard, Léonard Hussenot +4

cs.CL2026

Gemma 4 Technical Report

Gemma Team, Sherif El Abd, Vaibhav Aggarwal +320

cs.CL2024

Gemma: Open Models Based on Gemini Research and Technology

Gemma Team, Thomas Mesnard, Cassidy Hardin +105

cs.LG2022

Continuous Control with Action Quantization from Demonstrations

Robert Dadashi, Léonard Hussenot, Damien Vincent +4

cs.LG2020

CopyCAT: Taking Control of Neural Policies with Constant Attacks

Léonard Hussenot, Matthieu Geist, Olivier Pietquin

cs.LG2024

MusicRL: Aligning Music Generation to Human Preferences

Geoffrey Cideron, Sertan Girgin, Mauro Verzetti +11

cs.LG2021

Primal Wasserstein Imitation Learning

Robert Dadashi, Léonard Hussenot, Matthieu Geist +1

cs.LG2021

What Matters for Adversarial Imitation Learning?

Manu Orsini, Anton Raichuk, Léonard Hussenot +7

cs.LG2024

RecurrentGemma: Moving Past Transformers for Efficient Open Language Models

Aleksandar Botev, Soham De, Samuel L Smith +59

cs.CL2024

Gemma 2: Improving Open Language Models at a Practical Size

Gemma Team, Morgane Riviere, Shreya Pathak +195

cs.AI2026

MedGemma Technical Report

Andrew Sellergren, Sahar Kazemzadeh, Tiam Jaroensri +78

cs.CL2023

Factually Consistent Summarization via Reinforcement Learning with Textual Entailment Feedback

Paul Roit, Johan Ferret, Lior Shani +16

cs.LG2020

What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study

Marcin Andrychowicz, Anton Raichuk, Piotr Stańczyk +9

cs.LG2024

Conditional Language Policy: A General Framework for Steerable Multi-Objective Finetuning

Kaiwen Wang, Rahul Kidambi, Ryan Sullivan +17

cs.LG2024

BOND: Aligning LLMs with Best-of-N Distillation

Pier Giuseppe Sessa, Robert Dadashi, Léonard Hussenot +17