15 papers
Learning to Inject: Automated Prompt Injection via Reinforcement Learning
Xin Chen, Jie Zhang, Florian Tramèr
Prompt injection is a critical vulnerability in LLM agents, yet the strongest methods still rely on human red-teamers and hand-crafted prompts. Adapting automated jailbreak optimiz…
Position: Adversarial ML for LLMs Is Not Making Any Progress
Javier Rando, Jie Zhang, Nicholas Carlini +1
In the past decade, considerable research effort has been devoted to securing machine learning (ML) models that operate in adversarial settings. Yet, progress has been slow even fo…
Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning
Weitao Feng, Lixu Wang, Peizhuo Lv +5
As large language models (LLMs) continue to grow in capability, so do the risks of harmful misuse through fine-tuning. While most prior studies assume that attackers rely on superv…
MoAPT: Mixture of Adversarial Prompt Tuning for Vision-Language Models
Shiji Zhao, Qihui Zhu, Shukun Xiong +7
Large pre-trained Vision Language Models (VLMs) demonstrate excellent generalization capabilities but remain highly susceptible to adversarial examples, posing potential security r…
RealMath: A Continuous Benchmark for Evaluating Language Models on Research-Level Mathematics
Jie Zhang, Cezara Petrui, Kristina NikoliÄ +1
Existing benchmarks for evaluating mathematical reasoning in large language models (LLMs) rely primarily on competition problems, formal proofs, or artificially challenging questio…
Patronus: Safeguarding Text-to-Image Models against White-Box Adversaries
Xinfeng Li, Shengyuan Pang, Jialin Wu +5
Text-to-image (T2I) models, though exhibiting remarkable creativity in image generation, can be exploited to produce unsafe images. Existing safety measures, e.g., content moderati…