7 papers
Feature-Space Smoothing: Certified Robustness of Deep Representations
Song Xia, Meiwen Ding, Chenqi Kong +2
Modern deep learning models exhibit strong capabilities across diverse applications, yet remain vulnerable to malicious inputs that induce erroneous predictions via feature-space d…
Covert Visual Prompt Injection against Commercial Multimodal Large Language Models
Meiwen Ding, Song Xia, Chenqi Kong +1
Although multimodal large language models (MLLMs) are increasingly deployed in real-world applications, their instruction-following behavior leaves them vulnerable to prompt inject…
SAKED: Mitigating Hallucination in Large Vision-Language Models via Stability-Aware Knowledge Enhanced Decoding
Zhaoxu Li, Chenqi Kong, Peijun Bao +5
Hallucinations in Large Vision-Language Models (LVLMs) pose significant security and reliability risks in real-world applications. Inspired by the observation that humans are more…
From Pretrain to Pain: Adversarial Vulnerability of Video Foundation Models Without Task Knowledge
Hui Lu, Yi Yu, Song Xia +5
Large-scale Video Foundation Models (VFMs) has significantly advanced various video-related tasks, either through task-specific models or Multi-modal Large Language Models (MLLMs).…
Open-set Anomaly Segmentation in Complex Scenarios
Song Xia, Yi Yu, Henghui Ding +4
Precise segmentation of out-of-distribution (OoD) objects, herein referred to as anomalies, is crucial for the reliable deployment of semantic segmentation models in open-set, safe…
Transferable Adversarial Attacks on SAM and Its Downstream Models
Song Xia, Wenhan Yang, Yi Yu +4
The utilization of large foundational models has a dilemma: while fine-tuning downstream tasks from them holds promise for making use of the well-generalized knowledge in practical…