3 papers
cs.LG2026
Manifold Bandits: Bayesian Curriculum Learning over the Latent Geometry of Large Language Models
Darrien McKenzie, Nicklas Hansen, Xiaolong Wang
Reinforcement learning (RL) is a central approach for improving reasoning capabilities in large language models (LLMs), where training efficiency depends critically on how problems…
cs.LG2024
Maximum Entropy Hindsight Experience Replay
Douglas C. Crowder, Matthew L. Trappett, Darrien M. McKenzie +1
Hindsight experience replay (HER) is well-known to accelerate goal-based reinforcement learning (RL). While HER is generally applied to off-policy RL algorithms, we previously show…
cs.LG2024
Hindsight Experience Replay Accelerates Proximal Policy Optimization
Douglas C. Crowder, Darrien M. McKenzie, Matthew L. Trappett +1
Hindsight experience replay (HER) accelerates off-policy reinforcement learning algorithms for environments that emit sparse rewards by modifying the goal of the episode post-hoc t…