4 papers
Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors
Maheep Chaudhary, Fazl Barez
White-box monitoring is increasingly adopted as an auditing tool as Large Language Models (LLMs) are deployed in daily operations to ensure safe model behavior. However, white-box…
Punctuation and Predicates in Language Models
Sonakshi Chauhan, Maheep Chaudhary, Koby Choy +2
In this paper we explore where information is collected and how it is propagated throughout layers in large language models (LLMs). We begin by examining the surprising computation…
Studying Cross-cluster Modularity in Neural Networks
Satvik Golechha, Maheep Chaudhary, Joan Velja +2
An approach to improve neural network interpretability is via clusterability, i.e., splitting a model into disjoint clusters that can be studied independently. We define a measure…
Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability
Atticus Geiger, Duligur Ibeling, Amir Zur +8
Causal abstraction provides a theoretical foundation for mechanistic interpretability, the field concerned with providing intelligible algorithms that are faithful simplifications…