3 citations · 3 across the 11 of their papers we have counts for
3 papers · 1 filter
A Unified Theory of Sparse Dictionary Learning in Mechanistic Interpretability: Piecewise Biconvexity and Spurious Minima
Yiming Tang, Harshvardhan Saini, Zhaoqian Yao +6
As AI models achieve remarkable capabilities across diverse domains, understanding what representations they learn and how they encode concepts has become increasingly important fo…
Attribution Explanations for Deep Neural Networks: A Theoretical Perspective
Huiqi Deng, Hongbin Pei, Quanshi Zhang +1
Attribution explanation is a typical approach for explaining deep neural networks (DNNs), inferring an importance or contribution score for each input variable to the final output.…
AdaptiveK: Complexity-Driven Sparse Autoencoders for Interpretable Language Model Representations
Yifei Yao, Hanrong Zhang, Mengnan Du
Understanding the internal representations of large language models (LLMs) remains a central challenge for interpretability research. Sparse autoencoders (SAEs) offer a promising s…