most citedParaFuzz: An Interpretability-Driven Technique for Detecting Poisoned Samples in NLP

7 citations · 12 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CR2024

UNIT: Backdoor Mitigation via Automated Neural Distribution Tightening

Siyuan Cheng, Guangyu Shen, Kaiyuan Zhang +5

Deep neural networks (DNNs) have demonstrated effectiveness in various fields. However, DNNs are vulnerable to backdoor attacks, which inject a unique pattern, called trigger, into…

cs.NI2024

Laconic: Streamlined Load Balancers for SmartNICs

Tianyi Cui, Chenxingyu Zhao, Wei Zhang +2

Load balancers are pervasively used inside today's clouds to scalably distribute network requests across data center servers. Given the extensive use of load balancers and their as…

cs.AI20241 cited

Rapid Optimization for Jailbreaking LLMs via Subconscious Exploitation and Echopraxia

Guangyu Shen, Siyuan Cheng, Kaiyuan Zhang +6

Large Language Models (LLMs) have become prevalent across diverse sectors, transforming human life with their extraordinary reasoning and comprehension abilities. As they find incr…

cs.CR20237 cited

ParaFuzz: An Interpretability-Driven Technique for Detecting Poisoned Samples in NLP

Lu Yan, Zhuo Zhang, Guanhong Tao +4

Backdoor attacks have emerged as a prominent threat to natural language processing (NLP) models, where the presence of specific triggers in the input can lead poisoned models to mi…

cs.CV2023

Detecting Backdoors in Pre-trained Encoders

Shiwei Feng, Guanhong Tao, Siyuan Cheng +6

Self-supervised learning in computer vision trains on unlabeled data, such as images or (image, text) pairs, to obtain an image encoder that learns high-quality embeddings for inpu…

cs.CR20234 cited

BEAGLE: Forensics of Deep Learning Backdoor Attack for Better Defense

Siyuan Cheng, Guanhong Tao, Yingqi Liu +8

Deep Learning backdoor attacks have a threat model similar to traditional cyber attacks. Attack forensics, a critical counter-measure for traditional cyber attacks, is hence of imp…