3 papers
cs.LG2025
Stabilizing Policy Gradients for Sample-Efficient Reinforcement Learning in LLM Reasoning
Luckeciano C. Melo, Alessandro Abate, Yarin Gal
Reinforcement Learning, particularly through policy gradient methods, has played a central role in enabling reasoning capabilities of Large Language Models. However, the optimizati…
cs.LG2025
PAC Apprenticeship Learning with Bayesian Active Inverse Reinforcement Learning
Ondrej Bajgar, Dewi S. W. Gould, Jonathon Liu +3
As AI systems become increasingly autonomous, reliably aligning their decision-making with human preferences is essential. Inverse reinforcement learning (IRL) offers a promising a…
cs.LG2024
Temporal-Difference Variational Continual Learning
Luckeciano C. Melo, Alessandro Abate, Yarin Gal
Machine Learning models in real-world applications must continuously learn new tasks to adapt to shifts in the data-generating distribution. Yet, for Continual Learning (CL), model…