#adversarial evaluation
topicadversarial evaluation
2 papers · 1 filter
cs.CR2026
Attack Ensembles Expose a Safety-Utility Trade-off in Black-Box Guard Defenses Against Encoded VLM Jailbreaks
Haoyu Zhang, Zhuoxi Wang, Shibo Zheng +7
The paper proposes a guard‑agnostic recovery‑and‑decode module that transcribes encoded or visual text into plain language before applying existing safety classifiers for vision‑la…
#vision-language models#safety guards#jailbreak attacks#recovery decoding
cs.CL2026
Adversarial Pragmatics for AI Safety Evaluation: A Diagnostic Framework and Seed Benchmark for Language-Mediated Control
Brett Reynolds
The paper presents a benchmark and annotation protocol called adversarial pragmatics to evaluate language model safety when faced with instruction conflicts, embedded commands, and…
#adversarial evaluation#pragmatics#instruction conflict#embedded commands