24 citations · 53 across the 12 of their papers we have counts for
3 papers · 1 filter
Eight Methods to Evaluate Robust Unlearning in LLMs
Aengus Lynch, Phillip Guo, Aidan Ewart +2
Machine unlearning can be useful for removing harmful capabilities and memorized text from large language models (LLMs), but there are not yet standardized methods for rigorously e…
Cognitive Dissonance: Why Do Language Model Outputs Disagree with Internal Representations of Truthfulness?
Kevin Liu, Stephen Casper, Dylan Hadfield-Menell +1
Neural language models (LMs) can be used to evaluate the truth of factual statements in two ways: they can be either queried for statement probabilities, or probed for internal rep…
Explore, Establish, Exploit: Red Teaming Language Models from Scratch
Stephen Casper, Jason Lin, Joe Kwon +2
Deploying large language models (LMs) can pose hazards from harmful outputs such as toxic or false text. Prior work has introduced automated tools that elicit harmful outputs to id…