5 papers
Amortising Bayesian Experimental Design for Sequential Information Gathering in LLMs
Jakob Hartmann, James Harvey, Jhonathan Navott +5
Large language models (LLMs) exhibit strong reasoning and world-knowledge capabilities, yet often struggle to gather information effectively across the multi-turn interactions requ…
Stabilizing Policy Gradients for Sample-Efficient Reinforcement Learning in LLM Reasoning
Luckeciano C. Melo, Alessandro Abate, Yarin Gal
Reinforcement Learning, particularly through policy gradient methods, has played a central role in enabling reasoning capabilities of Large Language Models. However, the optimizati…
Temporal-Difference Variational Continual Learning
Luckeciano C. Melo, Alessandro Abate, Yarin Gal
Machine Learning models in real-world applications must continuously learn new tasks to adapt to shifts in the data-generating distribution. Yet, for Continual Learning (CL), model…
PAC Apprenticeship Learning with Bayesian Active Inverse Reinforcement Learning
Ondrej Bajgar, Dewi S. W. Gould, Jonathon Liu +3
As AI systems become increasingly autonomous, reliably aligning their decision-making with human preferences is essential. Inverse reinforcement learning (IRL) offers a promising a…
Deep Bayesian Active Learning for Preference Modeling in Large Language Models
Luckeciano C. Melo, Panagiotis Tigas, Alessandro Abate +1
Leveraging human preferences for steering the behavior of Large Language Models (LLMs) has demonstrated notable success in recent years. Nonetheless, data selection and labeling ar…