5 papers
Multi-Granular Node Pruning for Causal Circuit Discovery
Muhammad Umair Haider, Hammad Rizwan, Hassan Sajjad +1
Circuit discovery aims to identify minimal subnetworks that are responsible for specific behaviors in large language models (LLMs). Existing approaches primarily rely on iterative…
On the Persistent Effects of Lexicality in Large Language Models
Hammad Rizwan, Muhammad Umair Haider, Nishant Subramani +3
Representations extracted from large language models (LLMs) play an important role in many downstream applications. However, the structure of these representations is often influen…
Neurons Speak in Ranges: Breaking Free from Discrete Neuronal Attribution
Muhammad Umair Haider, Hammad Rizwan, Hassan Sajjad +2
Pervasive polysemanticity in large language models (LLMs) undermines discrete neuron-concept attribution, posing a significant challenge for model interpretation and control. We sy…
Resolving Lexical Bias in Model Editing
Hammad Rizwan, Domenic Rosati, Ga Wu +1
Model editing aims to modify the outputs of large language models after they are trained. Previous approaches have often involved direct alterations to model weights, which can res…
Instance-Level Difficulty: A Missing Perspective in Machine Unlearning
Hammad Rizwan, Mahtab Sarvmaili, Hassan Sajjad +1
Current research on deep machine unlearning primarily focuses on improving or evaluating the overall effectiveness of unlearning methods while overlooking the varying difficulty of…