5 papers
Compositional Planning with Jumpy World Models
Jesse Farebrother, Matteo Pirotta, Andrea Tirinzoni +3
The ability to plan with temporal abstractions is central to intelligent decision-making. Rather than reasoning over primitive actions, we study agents that compose pre-trained pol…
Convergence Theorems for Entropy-Regularized and Distributional Reinforcement Learning
Yash Jhaveri, Harley Wiltzer, Patrick Shafto +2
In the pursuit of finding an optimal policy, reinforcement learning (RL) methods generally ignore the properties of learned policies apart from their expected return. Thus, even wh…
Tapered Off-Policy REINFORCE: Stable and efficient reinforcement learning for LLMs
Nicolas Le Roux, Marc G. Bellemare, Jonathan Lebensold +7
We propose a new algorithm for fine-tuning large language models using reinforcement learning. Tapered Off-Policy REINFORCE (TOPR) uses an asymmetric, tapered variant of importance…
Action Gaps and Advantages in Continuous-Time Distributional Reinforcement Learning
Harley Wiltzer, Marc G. Bellemare, David Meger +2
When decisions are made at high frequency, traditional reinforcement learning (RL) methods struggle to accurately estimate action values. In turn, their performance is inconsistent…
Controlling Large Language Model Agents with Entropic Activation Steering
Nate Rahn, Pierluca D'Oro, Marc G. Bellemare
The rise of large language models (LLMs) has prompted increasing interest in their use as in-context learning agents. At the core of agentic behavior is the capacity for exploratio…