18 citations · 38 across the 5 of their papers we have counts for
5 papers
Can the Inference Logic of Large Language Models be Disentangled into Symbolic Concepts?
Wen Shen, Lei Cheng, Yuxiao Yang +2
In this paper, we explain the inference logic of large language models (LLMs) as a set of symbolic concepts. Many recent studies have discovered that traditional DNNs usually encod…
Understanding and Unifying Fourteen Attribution Methods with Taylor Interactions
Huiqi Deng, Na Zou, Mengnan Du +5
Various attribution methods have been developed to explain deep neural networks (DNNs) by inferring the attribution/importance/contribution score of each input variable to the fina…
Quantifying the Knowledge in a DNN to Explain Knowledge Distillation for Classification
Quanshi Zhang, Xu Cheng, Yilan Chen +1
Compared to traditional learning from scratch, knowledge distillation sometimes makes the DNN achieve superior performance. This paper provides a new perspective to explain the suc…
Proving Common Mechanisms Shared by Twelve Methods of Boosting Adversarial Transferability
Quanshi Zhang, Xin Wang, Jie Ren +4
Although many methods have been proposed to enhance the transferability of adversarial perturbations, these methods are designed in a heuristic manner, and the essential mechanism…
Growing Interpretable Part Graphs on ConvNets via Multi-Shot Learning
Quanshi Zhang, Ruiming Cao, Ying Nian Wu +1
This paper proposes a learning strategy that extracts object-part concepts from a pre-trained convolutional neural network (CNN), in an attempt to 1) explore explicit semantics hid…