2 papers
cs.CR2026
DE-FIVE: Detecting Malicious Image Prompts via Fourier Features and Image Vector Embeddings
Xingwei Zhong, Varun Sharma, Kar Wai Fok +1
Vision language models (VLMs) employ both visual and textual modalities to enable advanced vision-language inference. However, incorporating visual modalities expands the attack su…
cs.CR2025
Enhanced MLLM Black-Box Jailbreaking Attacks and Defenses
Xingwei Zhong, Kar Wai Fok, Vrizlynn L. L. Thing
Multimodal large language models (MLLMs) comprise of both visual and textual modalities to process vision language tasks. However, MLLMs are vulnerable to security-related issues,…