5 papers
CSO-LLM: Class Subspace Orthogonalization for Post-Training Backdoor Detection and Trigger Inversion in LLMs
Zhengxing Li, David J. Miller, Guangmingmei Yang +1
While post-training backdoor detection and trigger inversion schemes have been developed for AIs used e.g. for images, there is a paucity of such methods for LLMs. First, the LLM i…
A Novel Latent-Class Attack and its Detection by Class Subspace Orthogonalization
Guangmingmei Yang, David J. Miller, George Kesidis
Deep learning, which in general relies on voluminous amounts of training data, is vulnerable to data poisoning attacks, including error-generic attacks and backdoors (Trojans). In…
Improving the Sensitivity of Backdoor Detectors via Class Subspace Orthogonalization
Guangmingmei Yang, David J. Miller, George Kesidis
Most post-training backdoor detection methods rely on attacked models exhibiting extreme outlier detection statistics for the target class of an attack, compared to non-target clas…
Inverting Trojans in LLMs
Zhengxing Li, Guangmingmei Yang, Jayaram Raghuram +2
While effective backdoor detection and inversion schemes have been developed for AIs used e.g. for images, there are challenges in "porting" these methods to LLMs. First, the LLM i…
CEPA: Consensus Embedded Perturbation for Agnostic Detection and Inversion of Backdoors
Guangmingmei Yang, Xi Li, Hang Wang +2
A variety of defenses have been proposed against Trojans planted in (backdoor attacks on) deep neural network (DNN) classifiers. Backdoor-agnostic methods seek to reliably detect a…