4 papers
PoisonForge: Task-Level Targeted Poisoning Benchmark for Instruction-Tuned LLMs
Luze Sun, Anshuman Suri, Harsh Chaudhari +2
When practitioners fine-tune LLMs on unvetted datasets, an adversary can exploit the data supply chain through task-level poisoning: inserting a small number of crafted instruction…
Syntax- and Compilation-Preserving Evasion of LLM Vulnerability Detectors
Luze Sun, Alina Oprea, Eric Wong
LLM-based vulnerability detectors are increasingly deployed in CI/CD security gating, yet their resilience to evasion under syntax- and compilation-preserving edits remains poorly…
Benchmarking Misuse Mitigation Against Covert Adversaries
Davis Brown, Mahdi Sabbaghi, Luze Sun +4
Existing language model safety evaluations focus on overt attacks and low-stakes tasks. In reality, an attacker can easily subvert existing safeguards by requesting help on small,…
The Jailbreak Tax: How Useful are Your Jailbreak Outputs?
Kristina NikoliÄ, Luze Sun, Jie Zhang +1
Jailbreak attacks bypass the guardrails of large language models to produce harmful outputs. In this paper, we ask whether the model outputs produced by existing jailbreaks are act…