10 citations · 22 across the 20 of their papers we have counts for
11 papers · 1 filter
Defense Against LLM Backdoors using Critical Neuron Isolation Pruning
Yuxi Li, Zhibo Zhang, Kailong Wang +3
Large language models (LLMs) are vulnerable to backdoor attacks, where hidden triggers induce malicious outputs. Existing defenses generally fall into inference-time detection or t…
When Search Goes Wrong: Red-Teaming Web-Augmented Large Language Models
Haoran Ou, Kangjie Chen, Xingshuo Han +4
Large Language Models (LLMs) have been augmented with web search to overcome the limitations of the static knowledge boundary by accessing up-to-date information from the open Inte…
SwitchPatch: Physical Adversarial Attack Strategy with Switchable Adversarial Objectives
Hanrui Jiang, Yutong Wu, Shiyi Yao +5
Physical adversarial patch (PAP) attacks attach carefully crafted patches to physical objects to manipulate a deployed model. However, existing PAP attacks suffer from several limi…
SSD: A State-based Stealthy Backdoor Attack For Navigation System in UAV Route Planning
Zhaoxuan Wang, Yang Li, Jie Zhang +6
Unmanned aerial vehicles (UAVs) are increasingly employed to perform high-risk tasks that require minimal human intervention. However, UAVs face escalating cybersecurity threats, p…
PEFTGuard: Detecting Backdoor Attacks Against Parameter-Efficient Fine-Tuning
Zhen Sun, Tianshuo Cong, Yule Liu +5
Fine-tuning is an essential process to improve the performance of Large Language Models (LLMs) in specific domains, with Parameter-Efficient Fine-Tuning (PEFT) gaining popularity d…
VerifyML: Obliviously Checking Model Fairness Resilient to Malicious Model Holder
Guowen Xu, Xingshuo Han, Gelei Deng +5
In this paper, we present VerifyML, the first secure inference framework to check the fairness degree of a given Machine learning (ML) model. VerifyML is generic and is immune to a…