activity
20162023
most citedUnderstanding and Unifying Fourteen Attribution Methods with Taylor Interactions

18 citations · 38 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CL20231 cited

Can the Inference Logic of Large Language Models be Disentangled into Symbolic Concepts?

Wen Shen, Lei Cheng, Yuxiao Yang +2

In this paper, we explain the inference logic of large language models (LLMs) as a set of symbolic concepts. Many recent studies have discovered that traditional DNNs usually encod…

cs.LG202318 cited

Understanding and Unifying Fourteen Attribution Methods with Taylor Interactions

Huiqi Deng, Na Zou, Mengnan Du +5

Various attribution methods have been developed to explain deep neural networks (DNNs) by inferring the attribution/importance/contribution score of each input variable to the fina…

cs.LG20227 cited

Quantifying the Knowledge in a DNN to Explain Knowledge Distillation for Classification

Quanshi Zhang, Xu Cheng, Yilan Chen +1

Compared to traditional learning from scratch, knowledge distillation sometimes makes the DNN achieve superior performance. This paper provides a new perspective to explain the suc…

cs.LG20222 cited

Proving Common Mechanisms Shared by Twelve Methods of Boosting Adversarial Transferability

Quanshi Zhang, Xin Wang, Jie Ren +4

Although many methods have been proposed to enhance the transferability of adversarial perturbations, these methods are designed in a heuristic manner, and the essential mechanism…

cs.CV201610 cited

Growing Interpretable Part Graphs on ConvNets via Multi-Shot Learning

Quanshi Zhang, Ruiming Cao, Ying Nian Wu +1

This paper proposes a learning strategy that extracts object-part concepts from a pre-trained convolutional neural network (CNN), in an attempt to 1) explore explicit semantics hid…