5 papers
Multi-Granular Node Pruning for Causal Circuit Discovery
Muhammad Umair Haider, Hammad Rizwan, Hassan Sajjad +1
Circuit discovery aims to identify minimal subnetworks that are responsible for specific behaviors in large language models (LLMs). Existing approaches primarily rely on iterative…
On the Persistent Effects of Lexicality in Large Language Models
Hammad Rizwan, Muhammad Umair Haider, Nishant Subramani +3
Representations extracted from large language models (LLMs) play an important role in many downstream applications. However, the structure of these representations is often influen…
Neurons Speak in Ranges: Breaking Free from Discrete Neuronal Attribution
Muhammad Umair Haider, Hammad Rizwan, Hassan Sajjad +2
Pervasive polysemanticity in large language models (LLMs) undermines discrete neuron-concept attribution, posing a significant challenge for model interpretation and control. We sy…
Evaluating Sparse Autoencoders for Monosemantic Representation
Moghis Fereidouni, Muhammad Umair Haider, Peizhong Ju +1
A key barrier to interpreting large language models is polysemanticity, where neurons activate for multiple unrelated concepts. Sparse autoencoders (SAEs) have been proposed to mit…
Looking into Black Box Code Language Models
Muhammad Umair Haider, Umar Farooq, A. B. Siddique +1
Language Models (LMs) have shown their application for tasks pertinent to code and several code~LMs have been proposed recently. The majority of the studies in this direction only…