activity
20182025
most citedConstrained Optimization with Dynamic Bound-scaling for Effective NLPBackdoor Defense

11 citations · 38 across the 12 of their papers we have counts for

collaborators

8 papers

cs.CR2025

Temporal Logic-Based Multi-Vehicle Backdoor Attacks against Offline RL Agents in End-to-end Autonomous Driving

Xuan Chen, Shiwei Feng, Zikang Xiong +6

Assessing the safety of autonomous driving (AD) systems against security threats, particularly backdoor attacks, is a stepping stone for real-world deployment. However, existing wo…

cs.CR20226 cited

Backdoor Vulnerabilities in Normally Trained Deep Learning Models

Guanhong Tao, Zhenting Wang, Siyuan Cheng +7

We conduct a systematic study of backdoor vulnerabilities in normally trained Deep Learning models. They are as dangerous as backdoors injected by data poisoning because both can b…

cs.CL202211 cited

Constrained Optimization with Dynamic Bound-scaling for Effective NLPBackdoor Defense

Guangyu Shen, Yingqi Liu, Guanhong Tao +5

We develop a novel optimization method for NLPbackdoor inversion. We leverage a dynamically reducing temperature coefficient in the softmax function to provide changing loss landsc…

cs.LG20217 cited

EX-RAY: Distinguishing Injected Backdoor from Natural Features in Neural Networks by Examining Differential Feature Symmetry

Yingqi Liu, Guangyu Shen, Guanhong Tao +3

Backdoor attack injects malicious behavior to models such that inputs embedded with triggers are misclassified to a target label desired by the attacker. However, natural features…

cs.LG2021

Backdoor Scanning for Deep Neural Networks through K-Arm Optimization

Guangyu Shen, Yingqi Liu, Guanhong Tao +5

Back-door attack poses a severe threat to deep learning systems. It injects hidden malicious behaviors to a model such that any input stamped with a special pattern can trigger suc…

cs.LG2020

D-square-B: Deep Distribution Bound for Natural-looking Adversarial Attack

Qiuling Xu, Guanhong Tao, Xiangyu Zhang

We propose a novel technique that can generate natural-looking adversarial examples by bounding the variations induced for internal activation values in some deep layer(s), through…