most citedPAC-Bayesian Lifelong Learning For Multi-Armed Bandits

8 citations · 16 across the 9 of their papers we have counts for

collaborators

9 papers

cs.LG20223 cited

Information-Theoretic Safe Exploration with Gaussian Processes

Alessandro G. Bottero, Carlos E. Luis, Julia Vinogradska +2

We consider a sequential decision making task where we are not allowed to evaluate parameters that violate an a priori unknown (safety) constraint. A common approach is to place a…

cs.LG20224 cited

Structured Q-learning For Antibody Design

Alexander I. Cowen-Rivers, Philip John Gorinski, Aivar Sootla +5

Optimizing combinatorial structures is core to many real-world problems, such as those encountered in life sciences. For example, one of the crucial steps involved in antibody desi…

cs.LG2022

Self-supervised Sequential Information Bottleneck for Robust Exploration in Deep Reinforcement Learning

Bang You, Jingming Xie, Youping Chen +2

Effective exploration is critical for reinforcement learning agents in environments with sparse rewards or high-dimensional state-action spaces. Recent works based on state-visitat…

cs.LG2022

Revisiting Model-based Value Expansion

Daniel Palenicek, Michael Lutter, Jan Peters

Model-based value expansion methods promise to improve the quality of value function targets and, thereby, the effectiveness of value function learning. However, to date, these met…

cs.LG20221 cited

Dimensionality Reduction and Prioritized Exploration for Policy Search

Marius Memmel, Puze Liu, Davide Tateo +1

Black-box policy optimization is a class of reinforcement learning algorithms that explores and updates the policies at the parameter level. This class of algorithms is widely appl…

cs.LG2022

An Analysis of Measure-Valued Derivatives for Policy Gradients

Joao Carvalho, Jan Peters

Reinforcement learning methods for robotics are increasingly successful due to the constant development of better policy gradient techniques. A precise (low variance) and accurate…