papers

Publications (19)

cs.CR2025

MGC: A Compiler Framework Exploiting Compositional Blindness in Aligned LLMs for Malware Generation

Lu Yan, Zhuo Zhang, Xiangzhe Xu +5

Large language models (LLMs) have democratized software development, reducing the expertise barrier for programming complex applications. This accessibility extends to malicious so…

cs.CR2023

BEAGLE: Forensics of Deep Learning Backdoor Attack for Better Defense

Siyuan Cheng, Guanhong Tao, Yingqi Liu +8

Deep Learning backdoor attacks have a threat model similar to traditional cyber attacks. Attack forensics, a critical counter-measure for traditional cyber attacks, is hence of imp…

cs.LG2021

Backdoor Scanning for Deep Neural Networks through K-Arm Optimization

Guangyu Shen, Yingqi Liu, Guanhong Tao +5

Back-door attack poses a severe threat to deep learning systems. It injects hidden malicious behaviors to a model such that any input stamped with a special pattern can trigger suc…

cs.CV2024

LOTUS: Evasive and Resilient Backdoor Attacks through Sub-Partitioning

Siyuan Cheng, Guanhong Tao, Yingqi Liu +7

Backdoor attack poses a significant security threat to Deep Learning applications. Existing attacks are often not evasive to established backdoor detection techniques. This suscept…

cs.HC2026

Assistant or Actor? Student Trust, Control, and Delegation Regret When Using a General-Purpose AI Agent

Shiva Pochampally, Shengwei An, Yan Chen

When AI agents shift from answering questions to taking actions, users face a new problem: deciding what to delegate, to a system whose action space they cannot fully anticipate. W…

cs.CL2022

Constrained Optimization with Dynamic Bound-scaling for Effective NLPBackdoor Defense

Guangyu Shen, Yingqi Liu, Guanhong Tao +5

We develop a novel optimization method for NLPbackdoor inversion. We leverage a dynamically reducing temperature coefficient in the softmax function to provide changing loss landsc…

cs.LG2025

CENSOR: Defense Against Gradient Inversion via Orthogonal Subspace Bayesian Sampling

Kaiyuan Zhang, Siyuan Cheng, Guangyu Shen +5

Federated learning collaboratively trains a neural network on a global server, where each local client receives the current global model weights and sends back parameter updates (g…

cs.CR2022

DECK: Model Hardening for Defending Pervasive Backdoors

Guanhong Tao, Yingqi Liu, Siyuan Cheng +5

Pervasive backdoors are triggered by dynamic and pervasive input perturbations. They can be intentionally injected by attackers or naturally exist in normally trained models. They…

cs.AI2026

ManimAgent: Self-Evolving Multimodal Agents for Visual Education

Wenjia Jiang, Zongyuan Cai, Yuanhang Shao +7

Multi-round reflection lets agents built on large language models recover from failures within a single task, but each task remains an isolated episode: lessons learned across many…

cs.CR2023

FLIP: A Provable Defense Framework for Backdoor Mitigation in Federated Learning

Kaiyuan Zhang, Guanhong Tao, Qiuling Xu +8

Federated Learning (FL) is a distributed learning paradigm that enables different parties to train a model together for high quality and strong privacy protection. In this scenario…

cs.CR2021

Backdoor Attack through Frequency Domain

Tong Wang, Yuan Yao, Feng Xu +3

Backdoor attacks have been shown to be a serious threat against deep learning systems such as biometric authentication and autonomous driving. An effective backdoor attack could en…

cs.LG2026

Backdooring Masked Diffusion Language Models

Daniel Yiming Cao, Chengzhong Wang, Sheng-Yen Chou +3

Masked diffusion language models (MDLMs) are emerging as a compelling new paradigm for text generation, but their training-time security remains largely unexplored. Existing backdo…

cs.CR2025

Temporal Logic-Based Multi-Vehicle Backdoor Attacks against Offline RL Agents in End-to-end Autonomous Driving

Xuan Chen, Shiwei Feng, Zikang Xiong +6

Assessing the safety of autonomous driving (AD) systems against security threats, particularly backdoor attacks, is a stepping stone for real-world deployment. However, existing wo…

cs.CR2024

Elijah: Eliminating Backdoors Injected in Diffusion Models via Distribution Shift

Shengwei An, Sheng-Yen Chou, Kaiyuan Zhang +8

Diffusion models (DM) have become state-of-the-art generative models because of their capability to generate high-quality images from noises without adversarial training. However,…

cs.CR2022

Confidence Matters: Inspecting Backdoors in Deep Neural Networks via Distribution Transfer

Tong Wang, Yuan Yao, Feng Xu +3

Backdoor attacks have been shown to be a serious security threat against deep learning models, and detecting whether a given model has been backdoored becomes a crucial task. Exist…

cs.CR2024

UNIT: Backdoor Mitigation via Automated Neural Distribution Tightening

Siyuan Cheng, Guangyu Shen, Kaiyuan Zhang +5

Deep neural networks (DNNs) have demonstrated effectiveness in various fields. However, DNNs are vulnerable to backdoor attacks, which inject a unique pattern, called trigger, into…

cs.AI2024

Rapid Optimization for Jailbreaking LLMs via Subconscious Exploitation and Echopraxia

Guangyu Shen, Siyuan Cheng, Kaiyuan Zhang +6

Large Language Models (LLMs) have become prevalent across diverse sectors, transforming human life with their extraordinary reasoning and comprehension abilities. As they find incr…

cs.CR2022

Backdoor Vulnerabilities in Normally Trained Deep Learning Models

Guanhong Tao, Zhenting Wang, Siyuan Cheng +7

We conduct a systematic study of backdoor vulnerabilities in normally trained Deep Learning models. They are as dangerous as backdoors injected by data poisoning because both can b…

cs.CR2025

SOFT: Selective Data Obfuscation for Protecting LLM Fine-tuning against Membership Inference Attacks

Kaiyuan Zhang, Siyuan Cheng, Hanxi Guo +8

Large language models (LLMs) have achieved remarkable success and are widely adopted for diverse applications. However, fine-tuning these models often involves private or sensitive…