#jailbreak attacks

try —

5 papers match

cs.CR2026

Attack Ensembles Expose a Safety-Utility Trade-off in Black-Box Guard Defenses Against Encoded VLM Jailbreaks

Haoyu Zhang, Zhuoxi Wang, Shibo Zheng +7

The paper proposes a guard‑agnostic recovery‑and‑decode module that transcribes encoded or visual text into plain language before applying existing safety classifiers for vision‑la…

#vision-language models#safety guards#jailbreak attacks#recovery decoding
cs.SD2026

Prosody-driven Jailbreaks in Audio LLMs: A Controlled Study and Mechanistic Analysis

Jiachen Qian, Junyu Li

The paper investigates how variations in speech prosody, while keeping the transcript unchanged, can cause jailbreaks in audio-capable language models, introducing a new evaluation…

#audio language models#jailbreak attacks#prosody manipulation#safety evaluation
cs.CR2026

Borrowed Strength: Best-of-N Search over a Code EncodingBreaks Self-Check Jailbreak Defenses

Haoyu Zhang, Shibo Zheng, Xiangchen Guan +4

The paper demonstrates that a self‑check defense (SAGE) for language models can be bypassed by combining a code‑completion encoding attack with a best‑of‑N search, dramatically inc…

#jailbreak attacks#self‑check defense#code encoding#best‑of‑N search
cs.CL2026

Breaking Refusal in the First Half: A Mechanistic Study of the Prefill Jailbreak

Alex Kwon

The paper investigates why a short prefill phrase can disable the refusal behavior of aligned language models, pinpointing the failure to an early part of the model's response and…

#language model safety#jailbreak attacks#response‑site mechanisms#causal probing
cs.CL2026

HarDBench: A Benchmark for Draft-Based Co-Authoring Jailbreak Attacks for Safe Human-LLM Collaborative Writing

Euntae Kim, Soomin Han, Buru Chang

The paper introduces HarDBench, a benchmark that evaluates how vulnerable large language models are to jailbreak attacks when used as co-authors in draft-based writing, and propose…

#llm safety#jailbreak attacks#co-authoring#benchmark

One search, two signals: results blend meaning (embedding similarity, so papers that never use your words still surface) with keyword matches on titles, abstracts and summaries. Free, no sign-in needed.