2 citations · 2 across the 2 of their papers we have counts for
1 paper · 1 filter
Asael Sorensen, Charles Brock, David Chamberlain +3
Mechanistic interpretability seeks to make verifiable statements about the internal behavior of large language models (LLMs). Many interpretability techniques struggle to scale wit…