11 citations · 15 across the 3 of their papers we have counts for
3 papers
cs.LG2024★ 11 cited
Evaluating Frontier Models for Dangerous Capabilities
Mary Phuong, Matthew Aitchison, Elliot Catt +24
To understand the risks posed by a new AI system, we must understand what it can and cannot do. Building on prior work, we introduce a programme of new "dangerous capability" evalu…
cs.LG2023★ 1 cited
RLHF-Blender: A Configurable Interactive Interface for Learning from Diverse Human Feedback
Yannick Metz, David Lindner, Raphaël Baur +2
To use reinforcement learning from human feedback (RLHF) in practical applications, it is crucial to learn reward models from diverse sources of human feedback and to consider huma…
cs.LG2022★ 3 cited
Humans are not Boltzmann Distributions: Challenges and Opportunities for Modelling Human Feedback and Interaction in Reinforcement Learning
David Lindner, Mennatallah El-Assady
Reinforcement learning (RL) commonly assumes access to well-specified reward functions, which many practical applications do not provide. Instead, recently, more work has explored…