collaborators

8 papers

cs.CL2026

Large Language Models Generate Harmful Responses Using a Distinct Mechanism, Shared Across Harm Types

Hadas Orgad, Boyi Wei, Kaden Zheng +4

Large language models remain vulnerable to jailbreaks that elicit harmful responses, yet the mechanism behind harmful response generation is poorly understood. Here, we investigate…

cs.LG2026

Interpretability Can Be Actionable

Hadas Orgad, Fazl Barez, Tal Haklay +9

Interpretability aims to explain the behavior of deep neural networks. Despite rapid growth, there is mounting concern that much of this work has not translated into practical impa…

cs.CL2026

Hidden Failures in Robustness: Why Supervised Uncertainty Quantification Needs Better Evaluation

Joe Stacey, Hadas Orgad, Kentaro Inui +2

Recent work has shown that the hidden states of large language models contain signals useful for uncertainty estimation and hallucination detection, motivating a growing interest i…

cs.AI2026

Agents of Chaos

Natalie Shapira, Chris Wendler, Avery Yen +35

We report an exploratory red-teaming study of autonomous language-model-powered agents deployed in a live laboratory environment with persistent memory, email accounts, Discord acc…

cs.LG2025

MIB: A Mechanistic Interpretability Benchmark

Aaron Mueller, Atticus Geiger, Sarah Wiegreffe +20

How can we know whether new mechanistic interpretability methods achieve real improvements? In pursuit of lasting evaluation standards, we propose MIB, a Mechanistic Interpretabili…

cs.CL2025

LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations

Hadas Orgad, Michael Toker, Zorik Gekhman +4

Large language models (LLMs) often produce errors, including factual inaccuracies, biases, and reasoning failures, collectively referred to as "hallucinations". Recent studies have…