5 papers
Unraveling the Real Working Mechanism and Inherent Flaws of GAE: A Method for Interpreting Transformer Processes from an Economic Perspective
Yongjin Cui, Xiaohui Fan
We observe a phenomenon that current algorithmic research in the field of explainable artificial intelligence primarily pursues better performance on several proxy metrics. On the…
Generic Interpretation Approach for Transformer Models Incorporating Heterogenous Attention Structures
Yongjin Cui, Xiaohui Fan, Huajun Chen
Transformer has significantly propelled the development of artificial intelligence, and certainly the development of agents as well. We categorize attention structures of Transform…
Debunking Grad-ECLIP: A Comprehensive Study on Its Incorrectness and Fundamental Principles for Model Interpretation
Yongjin Cui, Xiaohui Fan
Grad-ECLIP is published at ICML 2024 and represents a new Transformer interpretation technical route (intermediate features-based). First, this paper demonstrates that the intermed…
Transformer Interpretability from Perspective of Attention and Gradient
Yongjin Cui, Xiaohui Fan, Huajun Chen
Although researchers' attention is more focused on the performance of Transformer models, the interpretation of Transformer can never be ignored. Gradient is widely utilized in Tra…
The Neglected Baseline in Model Interpretation
Yongjin Cui, Xiaohui Fan
We observe that existing model interpretation methods generally ignore the baseline, and such neglect often results in imprecise or even incorrect interpretation. In this paper, we…