From the 2 of 4 linked papers with an AI index.
4 papers
Whose Refusal Is It? The Unmeasured Contribution of Black-Box Multimodal Guardrails
Haoyu Zhang, Xiao Luo, Haowen Xu +4
A black-box guardrail is evaluated as though the safety number it earns were its own. It is not. A defended pipeline holds two components that can refuse (the guardrail, and the ta…
Decoy Images Amplify Caption-Mediated Defenses Against Encoded Jailbreaks
Haoyu Zhang, Xiangchen Guan, Shibo Zheng +2
We report a counter-intuitive interaction between image inputs and existing black-box defenses on Vision--Language Models (VLMs): pairing an encoded jailbreak prompt with an unrela…
Borrowed Strength: Best-of-N Search over a Code EncodingBreaks Self-Check Jailbreak Defenses
Haoyu Zhang, Shibo Zheng, Xiangchen Guan +4
The paper demonstrates that a self‑check defense (SAGE) for language models can be bypassed by combining a code‑completion encoding attack with a best‑of‑N search, dramatically inc…
Attack Ensembles Expose a Safety-Utility Trade-off in Black-Box Guard Defenses Against Encoded VLM Jailbreaks
Haoyu Zhang, Zhuoxi Wang, Shibo Zheng +7
The paper proposes a guard‑agnostic recovery‑and‑decode module that transcribes encoded or visual text into plain language before applying existing safety classifiers for vision‑la…