2 papers
cs.CL2022
Locally Aggregated Feature Attribution on Natural Language Model Understanding
Sheng Zhang, Jin Wang, Haitao Jiang +1
With the growing popularity of deep-learning models, model understanding becomes more important. Much effort has been devoted to demystify deep neural networks for better interpret…
cs.LG2020
Explaining Away Attacks Against Neural Networks
Sean Saito, Jin Wang
We investigate the problem of identifying adversarial attacks on image-based neural networks. We present intriguing experimental results showing significant discrepancies between t…