works on

From the 1 of 6 linked papers with an AI index.

activity
20242026
collaborators

6 papers

cs.CV2026

PersGuard: Preventing Malicious Personalization in Text-to-Image Diffusion Models via Model Backdoors

Xinwei Liu, Xiaojun Jia, Yuan Xun +2

The paper proposes PersGuard, a backdoor-based method that embeds protective triggers into pre‑trained text‑to‑image diffusion models so that unauthorized fine‑tuning on protected…

cs.CR2025

TeleAI-Safety: A comprehensive LLM jailbreaking benchmark towards attacks, defenses, and evaluations

Xiuyuan Chen, Jian Zhao, Yuxiang He +10

While the deployment of large language models (LLMs) in high-value industries continues to expand, the systematic assessment of their safety against jailbreak and prompt-based atta…

cs.CV2025

GeoShield: Safeguarding Geolocation Privacy from Vision-Language Models via Adversarial Perturbations

Xinwei Liu, Xiaojun Jia, Yuan Xun +2

Vision-Language Models (VLMs) such as GPT-4o now demonstrate a remarkable ability to infer users' locations from public shared images, posing a substantial risk to geoprivacy. Alth…

cs.AI2025

The Emotional Baby Is Truly Deadly: Does your Multimodal Large Reasoning Model Have Emotional Flattery towards Humans?

Yuan Xun, Xiaojun Jia, Xinwei Liu +1

We observe that MLRMs oriented toward human-centric service are highly susceptible to user emotional cues during the deep-thinking stage, often overriding safety protocols or built…

cs.CR2025

Robust Anti-Backdoor Instruction Tuning in LVLMs

Yuan Xun, Siyuan Liang, Xiaojun Jia +2

Large visual language models (LVLMs) have demonstrated excellent instruction-following capabilities, yet remain vulnerable to stealthy backdoor attacks when finetuned using contami…

cs.CV2024

CleanerCLIP: Fine-grained Counterfactual Semantic Augmentation for Backdoor Defense in Contrastive Learning

Yuan Xun, Siyuan Liang, Xiaojun Jia +2

Pre-trained large models for multimodal contrastive learning, such as CLIP, have been widely recognized in the industry as highly susceptible to data-poisoned backdoor attacks. Thi…