2 papers
cs.CR2026
Blind, Not Weak: A Best-of-Suite Safety-Utility Frontier for Recover-and-Reguard Defenses Against Encoded VLM Jailbreaks
Haoyu Zhang, Zhuoxi Wang, Shibo Zheng +9
Safety classifiers ("guards") are the dominant black-box defense for vision-language models, yet a guard judges an input's surface form, not its meaning: a harmful request re-encod…
cs.CR2026
Depth, Not Breadth: Best-of-N Jailbreaking Beyond Surface Noise
Haoyu Zhang, Shibo Zheng, Hanwen Liu +9
Best-of-N jailbreaking spends a query budget on surface variation, scrambling and recasing a request until one draw lands. We ask what a budget buys when its variance is moved into…