activity
20172024
most citedInformation-Theoretic Confidence Bounds for Reinforcement Learning

13 citations · 32 across the 9 of their papers we have counts for

collaborators
Showing cs.LGShow all

10 papers · 1 filter

cs.LG2024

RLHF and IIA: Perverse Incentives

Wanqiao Xu, Shi Dong, Xiuyuan Lu +3

Existing algorithms for reinforcement learning from human feedback (RLHF) can incentivize responses at odds with preferences because they are based on models that assume independen…

cs.LG2023

Approximate Thompson Sampling via Epistemic Neural Networks

Ian Osband, Zheng Wen, Seyed Mohammad Asghari +4

Thompson sampling (TS) is a popular heuristic for action selection, but it requires sampling from a posterior distribution. Unfortunately, this can become computationally intractab…

cs.LG2022★ 1 cited

Robustness of Epinets against Distributional Shifts

Xiuyuan Lu, Ian Osband, Seyed Mohammad Asghari +4

Recent work introduced the epinet as a new approach to uncertainty modeling in deep learning. An epinet is a small neural network added to traditional neural networks, which, toget…

cs.LG2022★ 6 cited

Ensembles for Uncertainty Estimation: Benefits of Prior Functions and Bootstrapping

Vikranth Dwaracherla, Zheng Wen, Ian Osband +3

In machine learning, an agent needs to estimate uncertainty to efficiently explore and adapt and to make effective decisions. A common approach to uncertainty estimation maintains…

cs.LG2022★ 2 cited

An Analysis of Ensemble Sampling

Chao Qin, Zheng Wen, Xiuyuan Lu +1

Ensemble sampling serves as a practical approximation to Thompson sampling when maintaining an exact posterior distribution over model parameters is computationally intractable. In…

cs.LG2021

The Neural Testbed: Evaluating Joint Predictions

Ian Osband, Zheng Wen, Seyed Mohammad Asghari +7

Predictive distributions quantify uncertainties ignored by point estimates. This paper introduces The Neural Testbed: an open-source benchmark for controlled and principled evaluat…