2 papers
cs.CL2026
How Much Do Circuits Tell Us? Measuring the Consistency and Specificity of Language Model Circuits
Michael Li, Nishant Subramani
The circuits framework in mechanistic interpretability aims to identify causally important sparse subgraphs of model components, typically evaluated by measuring necessity and suff…
cs.CL2026
Model Internal Sleuthing: Finding Lexical Identity and Inflectional Features in Modern Language Models
Michael Li, Nishant Subramani
Large transformer-based language models dominate modern NLP, yet our understanding of how they encode linguistic information relies primarily on studies of early models like BERT a…