1.8k citations · 2.6k across the 30 of their papers we have counts for
8 papers · 1 filter
PipelineRL: Faster On-policy Reinforcement Learning for Long Sequence Generation
Alexandre Piché, Ehsan Kamalloo, Rafael Pardinas +2
Reinforcement Learning (RL) is increasingly utilized to enhance the reasoning capabilities of Large Language Models (LLMs). However, effectively scaling these RL methods presents s…
RepoFusion: Training Code Models to Understand Your Repository
Disha Shrivastava, Denis Kocetkov, Harm de Vries +2
Despite the huge success of Large Language Models (LLMs) in coding assistants like GitHub Copilot, these models struggle to understand the context present in the repository (e.g.,…
Combating False Negatives in Adversarial Imitation Learning
Konrad Zolna, Chitwan Saharia, Leonard Boussioux +4
In adversarial imitation learning, a discriminator is trained to differentiate agent episodes from expert demonstrations representing the desired behavior. However, as the trained…
Automated curriculum generation for Policy Gradients from Demonstrations
Anirudh Srinivasan, Dzmitry Bahdanau, Maxime Chevalier-Boisvert +1
In this paper, we present a technique that improves the process of training an agent (using RL) for instruction following. We develop a training curriculum that uses a nominal numb…
Learning to Compute Word Embeddings On the Fly
Dzmitry Bahdanau, Tom Bosc, Stanisław Jastrzębski +3
Words in natural language follow a Zipfian distribution whereby some words are frequent but most are rare. Learning representations for words in the "long tail" of this distributio…
Sequence Tutor: Conservative Fine-Tuning of Sequence Generation Models with KL-control
Natasha Jaques, Shixiang Gu, Dzmitry Bahdanau +3
This paper proposes a general method for improving the structure and quality of sequences generated by a recurrent neural network (RNN), while maintaining information originally le…