3 papers
cs.CR2026
Refusing Everything Looks Safe: Restoring the Benign Arm to Encoded-Prompt Evaluation
Haoyu Zhang, Haowen Xu, Xiao Luo +7
Encoded-prompt attacks are evaluated almost entirely on their harmful arm: a benchmark sends obfuscated harmful requests and reports how often the model complied. A high refusal ra…
cs.CR2026
The Uncontrolled Variable: Vision-Language Model Refusal Responds to Image Presence in Ways Risk Cannot Explain
Haoyu Zhang, Yi Feng, Shibo Zheng +6
Vision-Language Model (VLM) safety is expected to depend on what a request asks for. We show that safety-aligned VLMs also key refusal on a property of a request's form: whether an…
cs.CR2026
0%, 45%, or 99%: A Guardrail's Own Share of the Refusals It Is Credited With
Haoyu Zhang, Xiangchen Guan, Xiao Luo +7
A defended pipeline's refusals have two producers: the guardrail bolted in front of the model, and the model's own alignment. Recovering the split costs nothing, because a guard bl…