1 citations · 2 across the 20 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States
Yilei Jiang, Xinyan Gao, Tianshuo Peng +4
The integration of additional modalities increases the susceptibility of large vision-language models (LVLMs) to safety risks, such as jailbreak attacks, compared to their language…
cs.CL2024
RapGuard: Safeguarding Multimodal Large Language Models via Rationale-aware Defensive Prompting
Yilei Jiang, Yingshui Tan, Xiangyu Yue
While Multimodal Large Language Models (MLLMs) have made remarkable progress in vision-language reasoning, they are also more susceptible to producing harmful content compared to m…