8 papers
Large Language Models Generate Harmful Responses Using a Distinct Mechanism, Shared Across Harm Types
Hadas Orgad, Boyi Wei, Kaden Zheng +4
Large language models remain vulnerable to jailbreaks that elicit harmful responses, yet the mechanism behind harmful response generation is poorly understood. Here, we investigate…
Interpretability Can Be Actionable
Hadas Orgad, Fazl Barez, Tal Haklay +9
Interpretability aims to explain the behavior of deep neural networks. Despite rapid growth, there is mounting concern that much of this work has not translated into practical impa…
Hidden Failures in Robustness: Why Supervised Uncertainty Quantification Needs Better Evaluation
Joe Stacey, Hadas Orgad, Kentaro Inui +2
Recent work has shown that the hidden states of large language models contain signals useful for uncertainty estimation and hallucination detection, motivating a growing interest i…
Agents of Chaos
Natalie Shapira, Chris Wendler, Avery Yen +35
We report an exploratory red-teaming study of autonomous language-model-powered agents deployed in a live laboratory environment with persistent memory, email accounts, Discord acc…
MIB: A Mechanistic Interpretability Benchmark
Aaron Mueller, Atticus Geiger, Sarah Wiegreffe +20
How can we know whether new mechanistic interpretability methods achieve real improvements? In pursuit of lasting evaluation standards, we propose MIB, a Mechanistic Interpretabili…
LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
Hadas Orgad, Michael Toker, Zorik Gekhman +4
Large language models (LLMs) often produce errors, including factual inaccuracies, biases, and reasoning failures, collectively referred to as "hallucinations". Recent studies have…