7 papers
TrapSuffix: Proactive Defense Against Adversarial Suffixes in Jailbreaking
Mengyao Du, Han Fang, Haokai Ma +4
Suffix-based jailbreak attacks append an adversarial suffix, i.e., a short token sequence, to steer aligned LLMs into unsafe outputs. Since suffixes are free-form text, they admit…
Label Forensics: Interpreting Hard Labels in Black-Box Text Classifier
Mengyao Du, Gang Yang, Han Fang +2
The widespread adoption of natural language processing techniques has led to an unprecedented growth of text classifiers across the modern web. Yet many of these models circulate w…
Removal Attack and Defense on AI-generated Content Latent-based Watermarking
De Zhang Lee, Han Fang, Hanyi Wang +1
Digital watermarks can be embedded into AI-generated content (AIGC) by initializing the generation process with starting points sampled from a secret distribution. When combined wi…
SafeMobile: Chain-level Jailbreak Detection and Automated Evaluation for Multimodal Mobile Agents
Siyuan Liang, Tianmeng Fang, Zhe Liu +5
With the wide application of multimodal foundation models in intelligent agent systems, scenarios such as mobile device control, intelligent assistant interaction, and multimodal t…
Poison Once, Control Anywhere: Clean-Text Visual Backdoors in VLM-based Mobile Agents
Xuan Wang, Siyuan Liang, Zhe Liu +5
Mobile agents powered by vision-language models (VLMs) are increasingly adopted for tasks such as UI automation and camera-based assistance. These agents are typically fine-tuned u…
Lie Detector: Unified Backdoor Detection via Cross-Examination Framework
Xuan Wang, Siyuan Liang, Dongping Liao +6
Institutions with limited data and computing resources often outsource model training to third-party providers in a semi-honest setting, assuming adherence to prescribed training p…