Publications (39)
ShiQ: Bringing back Bellman to LLMs
Pierre Clavier, Nathan Grinsztajn, Raphael Avalos +8
Mastering the Game of Stratego with Model-Free Multiagent Reinforcement Learning
Julien Perolat, Bart de Vylder, Daniel Hennes +31
On the role of population heterogeneity in emergent communication
Mathieu Rita, Florian Strub, Jean-Bastien Grill +2
Broaden Your Views for Self-Supervised Video Learning
Adrià Recasens, Pauline Luc, Jean-Baptiste Alayrac +11
Command A: An Enterprise-Ready Large Language Model
Team Cohere, :, Aakanksha +227
Countering Reward Over-optimization in LLM with Demonstration-Guided Reinforcement Learning
Mathieu Rita, Florian Strub, Rahma Chaabouni +3
Contrastive Policy Gradient: Aligning LLMs on sequence-level scores in a supervised-friendly fashion
Yannis Flet-Berliac, Nathan Grinsztajn, Florian Strub +8
Learning Natural Language Generation from Scratch
Alice Martin Donati, Guillaume Quispe, Charles Ollion +3
Visual Reasoning with Multi-hop Feature Modulation
Florian Strub, Mathieu Seurin, Ethan Perez +5
HIGhER : Improving instruction following with Hindsight Generation for Experience Replay
Geoffrey Cideron, Mathieu Seurin, Florian Strub +1
The Edge of Orthogonality: A Simple View of What Makes BYOL Tick
Pierre H. Richemond, Allison Tam, Yunhao Tang +3
A Machine of Few Words -- Interactive Speaker Recognition with Reinforcement Learning
Mathieu Seurin, Florian Strub, Philippe Preux +1
Deep Reinforcement Learning and the Deadly Triad
Hado van Hasselt, Yotam Doron, Florian Strub +3
Correction of Electron Back-scattered Diffraction datasets using an evolutionary algorithm
Florian Strub, Marie-Agathe Charpagne, Tresa M. Pollock
World Modelling Improves Language Model Agents
Shangmin Guo, Omar Darwiche Domingues, Raphaël Avalos +2
Emergent Communication: Generalization and Overfitting in Lewis Games
Mathieu Rita, Corentin Tallec, Paul Michel +4
Hybrid Recommender System based on Autoencoders
Florian Strub, Romaric Gaudel, Jérémie Mary
Hybrid Collaborative Filtering with Autoencoders
Florian Strub, Jeremie Mary, Romaric Gaudel
GuessWhat?! Visual object discovery through multi-modal dialogue
Harm de Vries, Florian Strub, Sarath Chandar +3
Don't Do What Doesn't Matter: Intrinsic Motivation with Action Usefulness
Mathieu Seurin, Florian Strub, Philippe Preux +1
End-to-end optimization of goal-driven and visually grounded dialogue systems
Florian Strub, Harm de Vries, Jeremie Mary +3
Modulating early visual processing by language
Harm de Vries, Florian Strub, Jérémie Mary +3
Language Evolution with Deep Learning
Mathieu Rita, Paul Michel, Rahma Chaabouni +3
The Monte Carlo Transformer: a stochastic self-attention model for sequence prediction
Alice Martin, Charles Ollion, Florian Strub +2
Learning Nash Equilibrium for General-Sum Markov Games from Batch Data
Julien Pérolat, Florian Strub, Bilal Piot +1
Over-communicate no more: Situated RL agents learn concise communication protocols
Aleksandra Kalinowska, Elnaz Davoodi, Florian Strub +5
Aya Expanse: Combining Research Breakthroughs for a New Multilingual Frontier
John Dang, Shivalika Singh, Daniel D'souza +42
Developing, Evaluating and Scaling Learning Agents in Multi-Agent Environments
Ian Gemp, Thomas Anthony, Yoram Bachrach +24
Supervised Seeded Iterated Learning for Interactive Language Learning
Yuchen Lu, Soumye Singhal, Florian Strub +2
Bootstrap your own latent: A new approach to self-supervised Learning
Jean-Bastien Grill, Florian Strub, Florent Altché +11
Learning Visual Reasoning Without Strong Priors
Ethan Perez, Harm de Vries, Florian Strub +2
Averaging log-likelihoods in direct alignment
Nathan Grinsztajn, Yannis Flet-Berliac, Mohammad Gheshlaghi Azar +8
Accurate reconstruction of EBSD datasets by a multimodal data approach using an evolutionary algorithm
Marie-Agathe Charpagne, Florian Strub, Tresa M. Pollock
SemPPL: Predicting pseudo-labels for better contrastive representations
Matko Bošnjak, Pierre H. Richemond, Nenad Tomasev +7
HoME: a Household Multimodal Environment
Simon Brodeur, Ethan Perez, Ankesh Anand +6
BYOL works even without batch statistics
Pierre H. Richemond, Jean-Bastien Grill, Florent Altché +8
Language Model Alignment with Elastic Reset
Michael Noukhovitch, Samuel Lavoie, Florian Strub +1
Countering Language Drift with Seeded Iterated Learning
Yuchen Lu, Soumye Singhal, Florian Strub +2
FiLM: Visual Reasoning with a General Conditioning Layer
Ethan Perez, Florian Strub, Harm de Vries +2