activity
20222024
most citedEvaluating Frontier Models for Dangerous Capabilities

11 citations · 23 across the 8 of their papers we have counts for

collaborators
Showing cs.LGShow all

5 papers · 1 filter

cs.LG2024

Mapping out the Space of Human Feedback for Reinforcement Learning: A Conceptual Framework

Yannick Metz, David Lindner, Raphaël Baur +1

Reinforcement Learning from Human feedback (RLHF) has become a powerful tool to fine-tune or train agentic machine learning models. Similar to how humans interact in social context…

cs.LG20246 cited

On scalable oversight with weak LLMs judging strong LLMs

Zachary Kenton, Noah Y. Siegel, János Kramár +8

Scalable oversight protocols aim to enable humans to accurately supervise superhuman AI. In this paper we study debate, where two AI's compete to convince a judge; consultancy, whe…

cs.LG202411 cited

Evaluating Frontier Models for Dangerous Capabilities

Mary Phuong, Matthew Aitchison, Elliot Catt +24

To understand the risks posed by a new AI system, we must understand what it can and cannot do. Building on prior work, we introduce a programme of new "dangerous capability" evalu…

cs.LG20231 cited

RLHF-Blender: A Configurable Interactive Interface for Learning from Diverse Human Feedback

Yannick Metz, David Lindner, Raphaël Baur +2

To use reinforcement learning from human feedback (RLHF) in practical applications, it is crucial to learn reward models from diverse sources of human feedback and to consider huma…

cs.LG20223 cited

Humans are not Boltzmann Distributions: Challenges and Opportunities for Modelling Human Feedback and Interaction in Reinforcement Learning

David Lindner, Mennatallah El-Assady

Reinforcement learning (RL) commonly assumes access to well-specified reward functions, which many practical applications do not provide. Instead, recently, more work has explored…