From the 1 of 6 linked papers with an AI index.
6 papers
PersGuard: Preventing Malicious Personalization in Text-to-Image Diffusion Models via Model Backdoors
Xinwei Liu, Xiaojun Jia, Yuan Xun +2
The paper proposes PersGuard, a backdoor-based method that embeds protective triggers into pre‑trained text‑to‑image diffusion models so that unauthorized fine‑tuning on protected…
TeleAI-Safety: A comprehensive LLM jailbreaking benchmark towards attacks, defenses, and evaluations
Xiuyuan Chen, Jian Zhao, Yuxiang He +10
While the deployment of large language models (LLMs) in high-value industries continues to expand, the systematic assessment of their safety against jailbreak and prompt-based atta…
GeoShield: Safeguarding Geolocation Privacy from Vision-Language Models via Adversarial Perturbations
Xinwei Liu, Xiaojun Jia, Yuan Xun +2
Vision-Language Models (VLMs) such as GPT-4o now demonstrate a remarkable ability to infer users' locations from public shared images, posing a substantial risk to geoprivacy. Alth…
The Emotional Baby Is Truly Deadly: Does your Multimodal Large Reasoning Model Have Emotional Flattery towards Humans?
Yuan Xun, Xiaojun Jia, Xinwei Liu +1
We observe that MLRMs oriented toward human-centric service are highly susceptible to user emotional cues during the deep-thinking stage, often overriding safety protocols or built…
Robust Anti-Backdoor Instruction Tuning in LVLMs
Yuan Xun, Siyuan Liang, Xiaojun Jia +2
Large visual language models (LVLMs) have demonstrated excellent instruction-following capabilities, yet remain vulnerable to stealthy backdoor attacks when finetuned using contami…
CleanerCLIP: Fine-grained Counterfactual Semantic Augmentation for Backdoor Defense in Contrastive Learning
Yuan Xun, Siyuan Liang, Xiaojun Jia +2
Pre-trained large models for multimodal contrastive learning, such as CLIP, have been widely recognized in the industry as highly susceptible to data-poisoned backdoor attacks. Thi…