1 paper
Asael Sorensen, Charles Brock, David Chamberlain +3
Mechanistic interpretability seeks to make verifiable statements about the internal behavior of large language models (LLMs). Many interpretability techniques struggle to scale wit…