3 citations · 4 across the 11 of their papers we have counts for
1 paper · 1 filter
Zixuan Xu, Tiancheng He, Huahui Yi +7
Vision-language models remain susceptible to multimodal jailbreaks and over-refusal because safety hinges on both visual evidence and user intent, while many alignment pipelines su…