3 papers
cs.CL2026
How Much Do Circuits Tell Us? Measuring the Consistency and Specificity of Language Model Circuits
Michael Li, Nishant Subramani
The circuits framework in mechanistic interpretability aims to identify sparse subgraphs of model components that are causally responsible for a behavior, typically evaluated by me…
cs.CL2025
Model Internal Sleuthing: Finding Lexical Identity and Inflectional Features in Modern Language Models
Michael Li, Nishant Subramani
Large transformer-based language models dominate modern NLP, yet our understanding of how they encode linguistic information relies primarily on studies of early models like BERT a…
cs.CL2024
View From Above: A Framework for Evaluating Distribution Shifts in Model Behavior
Tanush Chopra, Michael Li, Jacob Haimes
When large language models (LLMs) are asked to perform certain tasks, how can we be sure that their learned representations align with reality? We propose a domain-agnostic framewo…