collaborators

7 papers

cs.LG2026

Feature-Space Smoothing: Certified Robustness of Deep Representations

Song Xia, Meiwen Ding, Chenqi Kong +2

Modern deep learning models exhibit strong capabilities across diverse applications, yet remain vulnerable to malicious inputs that induce erroneous predictions via feature-space d…

cs.CV2026

Covert Visual Prompt Injection against Commercial Multimodal Large Language Models

Meiwen Ding, Song Xia, Chenqi Kong +1

Although multimodal large language models (MLLMs) are increasingly deployed in real-world applications, their instruction-following behavior leaves them vulnerable to prompt inject…

cs.CV2026

SAKED: Mitigating Hallucination in Large Vision-Language Models via Stability-Aware Knowledge Enhanced Decoding

Zhaoxu Li, Chenqi Kong, Peijun Bao +5

Hallucinations in Large Vision-Language Models (LVLMs) pose significant security and reliability risks in real-world applications. Inspired by the observation that humans are more…

cs.CV2025

From Pretrain to Pain: Adversarial Vulnerability of Video Foundation Models Without Task Knowledge

Hui Lu, Yi Yu, Song Xia +5

Large-scale Video Foundation Models (VFMs) has significantly advanced various video-related tasks, either through task-specific models or Multi-modal Large Language Models (MLLMs).…

cs.CV2025

Open-set Anomaly Segmentation in Complex Scenarios

Song Xia, Yi Yu, Henghui Ding +4

Precise segmentation of out-of-distribution (OoD) objects, herein referred to as anomalies, is crucial for the reliable deployment of semantic segmentation models in open-set, safe…

cs.LG2025

Transferable Adversarial Attacks on SAM and Its Downstream Models

Song Xia, Wenhan Yang, Yi Yu +4

The utilization of large foundational models has a dilemma: while fine-tuning downstream tasks from them holds promise for making use of the well-generalized knowledge in practical…