activity
20242026
collaborators

5 papers

cs.LG2026

Compositional Planning with Jumpy World Models

Jesse Farebrother, Matteo Pirotta, Andrea Tirinzoni +3

The ability to plan with temporal abstractions is central to intelligent decision-making. Rather than reasoning over primitive actions, we study agents that compose pre-trained pol…

cs.LG2025

Convergence Theorems for Entropy-Regularized and Distributional Reinforcement Learning

Yash Jhaveri, Harley Wiltzer, Patrick Shafto +2

In the pursuit of finding an optimal policy, reinforcement learning (RL) methods generally ignore the properties of learned policies apart from their expected return. Thus, even wh…

cs.LG2025

Tapered Off-Policy REINFORCE: Stable and efficient reinforcement learning for LLMs

Nicolas Le Roux, Marc G. Bellemare, Jonathan Lebensold +7

We propose a new algorithm for fine-tuning large language models using reinforcement learning. Tapered Off-Policy REINFORCE (TOPR) uses an asymmetric, tapered variant of importance…

cs.LG2024

Action Gaps and Advantages in Continuous-Time Distributional Reinforcement Learning

Harley Wiltzer, Marc G. Bellemare, David Meger +2

When decisions are made at high frequency, traditional reinforcement learning (RL) methods struggle to accurately estimate action values. In turn, their performance is inconsistent…

cs.CL2024

Controlling Large Language Model Agents with Entropic Activation Steering

Nate Rahn, Pierluca D'Oro, Marc G. Bellemare

The rise of large language models (LLMs) has prompted increasing interest in their use as in-context learning agents. At the core of agentic behavior is the capacity for exploratio…