4 citations · 4 across the 3 of their papers we have counts for
9 papers
Large Language Models Generate Harmful Responses Using a Distinct Mechanism, Shared Across Harm Types
Hadas Orgad, Boyi Wei, Kaden Zheng +4
Large language models remain vulnerable to jailbreaks that elicit harmful responses, yet the mechanism behind harmful response generation is poorly understood. Here, we investigate…
Tensor Product Representation Probes Reveal Shared Structure Across Linear Directions
Andrew Lee, Fernanda Viégas, Martin Wattenberg
While researchers are finding concepts represented as linear directions in language models, a bag of linear directions fails to capture relational structure. To better understand t…
Into the Rabbit Hull: From Task-Relevant Concepts in DINO to Minkowski Geometry
Thomas Fel, Binxu Wang, Michael A. Lepori +8
DINOv2 is routinely deployed to recognize objects, scenes, and actions; yet the nature of what it perceives remains unknown. As a working baseline, we adopt the Linear Representati…
Beyond the Rosetta Stone: Unification Forces in Generalization Dynamics
Carter Blum, Katja Filippova, Ann Yuan +8
Large language models (LLMs) struggle with cross-lingual knowledge transfer: they hallucinate when asked in one language about facts expressed in a different language during traini…
Does visualization help AI understand data?
Victoria R. Li, Johnathan Sun, Martin Wattenberg
Charts and graphs help people analyze data, but can they also be useful to AI systems? To investigate this question, we perform a series of experiments with two commercial vision-l…
Can Interpretation Predict Behavior on Unseen Data?
Victoria R. Li, Jenny Kaufmann, Martin Wattenberg +3
Interpretability research often predicts model responses to targeted mechanistic interventions. But can we predict responses to unseen input data? We propose and demonstrate this a…