Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
How Much Do Circuits Tell Us? Measuring the Consistency and Specificity of Language Model Circuits
Michael Li, Nishant Subramani
The circuits framework in mechanistic interpretability aims to identify causally important sparse subgraphs of model components, typically evaluated by measuring necessity and suff…
cs.CL2026
Model Internal Sleuthing: Finding Lexical Identity and Inflectional Features in Modern Language Models
Michael Li, Nishant Subramani
Large transformer-based language models dominate modern NLP, yet our understanding of how they encode linguistic information relies primarily on studies of early models like BERT a…
cs.CL2024
View From Above: A Framework for Evaluating Distribution Shifts in Model Behavior
Tanush Chopra, Michael Li, Jacob Haimes
When large language models (LLMs) are asked to perform certain tasks, how can we be sure that their learned representations align with reality? We propose a domain-agnostic framewo…