1 paper
Qing Li, Jiahui Geng, Zongxiong Chen +3
Vision-language models (VLMs) demonstrate strong multimodal capabilities but have been found to be more susceptible to generating harmful content compared to their backbone large l…