1 paper
Haoyu Zhang, Zhuoxi Wang, Shibo Zheng +7
Safety classifiers ("guards") are the dominant black-box defense for vision-language models, yet a guard judges an input's surface form, not its meaning: a harmful request re-encod…