1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.CL2024
Whispering Experts: Neural Interventions for Toxicity Mitigation in Language Models
Xavier Suau, Pieter Delobelle, Katherine Metcalf +4
An important issue with Large Language Models (LLMs) is their undesired ability to generate toxic language. In this work, we show that the neurons responsible for toxicity can be d…
cs.LG2024
Hindsight PRIORs for Reward Learning from Human Preferences
Mudit Verma, Katherine Metcalf
Preference based Reinforcement Learning (PbRL) removes the need to hand specify a reward function by learning a reward from preference feedback over policy behaviors. Current appro…
cs.AI2024★ 1 cited
Sample-Efficient Preference-based Reinforcement Learning with Dynamics Aware Rewards
Katherine Metcalf, Miguel Sarabia, Natalie Mackraz +1
Preference-based reinforcement learning (PbRL) aligns a robot behavior with human preferences via a reward function learned from binary feedback over agent behaviors. We show that…