7 papers
TIGA: Trajectory-Injected Generative Attack against Black-box AIGC Detectors
Xia Du, Zhuosen Bao, Zheng Lin +6
Recent diffusion models have achieved remarkable realism in facial image synthesis, posing growing challenges to artificial intelligence-generated content (AIGC) forensic detectors…
When Backdoors Meet Partial Observability: Attacking Real-World Reinforcement Learning
Tairan Huang, Qingqing Ye, Yulin Jin +4
Backdoor attacks can cause reinforcement learning (RL) policies to behave normally under clean inputs while executing malicious behaviors when triggers are present. Existing RL bac…
LLM-Agnostic Semantic Representation Attack
Jiawei Lian, Jianhong Pan, Lefan Wang +4
Large Language Models (LLMs) increasingly employ alignment techniques to prevent harmful outputs. Despite these safeguards, attackers can circumvent them by crafting adversarial pr…
Semantic Representation Attack against Aligned Large Language Models
Jiawei Lian, Jianhong Pan, Lefan Wang +3
Large Language Models (LLMs) increasingly employ alignment techniques to prevent harmful outputs. Despite these safeguards, attackers can circumvent them by crafting prompts that i…
Attack Anything: Blind DNNs via Universal Background Adversarial Attack
Jiawei Lian, Shaohui Mei, Xiaofei Wang +5
It has been widely substantiated that deep neural networks (DNNs) are susceptible and vulnerable to adversarial perturbations. Existing studies mainly focus on performing attacks b…
Revealing the Intrinsic Ethical Vulnerability of Aligned Large Language Models
Jiawei Lian, Jianhong Pan, Lefan Wang +3
Large language models (LLMs) are foundational explorations to artificial general intelligence, yet their alignment with human values via instruction tuning and preference learning…