11 citations · 18 across the 6 of their papers we have counts for
7 papers
Pathologies of Pre-trained Language Models in Few-shot Fine-tuning
Hanjie Chen, Guoqing Zheng, Ahmed Hassan Awadallah +1
Although adapting pre-trained language models with few examples has shown promising performance on text classification, there is a lack of understanding of where the performance ga…
Adversarial Training for Improving Model Robustness? Look at Both Prediction and Interpretation
Hanjie Chen, Yangfeng Ji
Neural language models show vulnerability to adversarial examples which are semantically similar to their original counterparts with a few words replaced by their synonyms. A commo…
Perturbing Inputs for Fragile Interpretations in Deep Natural Language Processing
Sanchit Sinha, Hanjie Chen, Arshdeep Sekhon +2
Interpretability methods like Integrated Gradient and LIME are popular choices for explaining natural language model predictions with relative word importance scores. These interpr…
Explaining Neural Network Predictions on Sentence Pairs via Learning Word-Group Masks
Hanjie Chen, Song Feng, Jatin Ganhotra +4
Explaining neural network models is important for increasing their trustworthiness in real-world applications. Most existing methods generate post-hoc explanations for neural netwo…
Learning Variational Word Masks to Improve the Interpretability of Neural Text Classifiers
Hanjie Chen, Yangfeng Ji
To build an interpretable neural text classifier, most of the prior work has focused on designing inherently interpretable models or finding faithful explanations. A new line of wo…
Generating Hierarchical Explanations on Text Classification via Feature Interaction Detection
Hanjie Chen, Guangtao Zheng, Yangfeng Ji
Generating explanations for neural networks has become crucial for their applications in real-world with respect to reliability and trustworthiness. In natural language processing,…