5 papers
Q-based Variational Inverse Reinforcement Learning
Ondrej Bajgar, Peter Tisnikar, Alessandro Abate +2
The development of safe and beneficial AI requires that systems can learn and act in accordance with human preferences. However, explicitly specifying these preferences by hand is…
Constrained Bayesian Optimisation with Multiple Information Sources
Hauke Maathuis, Roeland De Breuker, Saullo Castro +1
Bayesian Optimisation (BO) under unknown constraints is particularly challenging when feasible regions are small. In such settings, existing methods that typically rely solely on e…
RIZZ: Routing Interactions to Near Zero-Interference Zones for Continual Adaptation of Black-Box Agents
Sonali Goel, Pranav Vaidhyanathan, Lucas Schorling +2
Large language models are increasingly deployed as long-lived agents that must adapt across users, tasks, domains, modalities, and feedback regimes without access to model weights.…
Canonical Regularisation of Wide Feature-Learning Neural Networks
George Whittle, Pranav Vaidhyanathan, Juliusz Ziomek +2
Wide neural networks in the feature-learning regime drive modern deep learning, and yet they remain far less studied than their kernel-regime counterparts. We consider a critical y…
Fully Offline Reinforcement Learning
Mattie Fellows, Clarisse Wibault, Uljad Berdica +3
Offline RL (ORL) promises safe and sample-efficient deployment but existing methods rely on undocumented online interactions for hyperparameter tuning and lack reliable fully offli…