From the 1 of 12 linked papers with an AI index.
12 papers
UniNDM: A Unified Noise-driven Detection and Mitigation Framework Against Sexual Content in Text-to-Image Generation
Yao Huang, Yitong Sun, Huanran Chen +8
Despite the impressive generative capabilities of text-to-image diffusion models, they remain vulnerable to implicit sexual prompts, where subtle cues disguised as benign terms or…
PersGuard: Preventing Malicious Personalization in Text-to-Image Diffusion Models via Model Backdoors
Xinwei Liu, Xiaojun Jia, Yuan Xun +2
The paper proposes PersGuard, a backdoor-based method that embeds protective triggers into pre‑trained text‑to‑image diffusion models so that unauthorized fine‑tuning on protected…
SkillJect: Effectively Automating Skill-Based Prompt Injection for Skill-Enabled Agents
Xiaojun Jia, Jie Liao, Simeng Qin +5
Agent skills extend LLM agents with task-specific instructions, executable scripts, and auxiliary resources, improving reusability but creating a new supply-chain attack surface. A…
Need to Know: Contextual-Integrity-Grounded Query Rewriting for Privacy-Conscious LLM Delegation
Xinyue Huang, Xiaochun Cao, Wenyuan Yang
As LLMs become increasingly woven into everyday workflows, user queries sent to cloud hosted LLMs routinely mix task-essential content with task non-essential sensitive disclosures…
Reasoning as an Attack Surface: Adaptive Evolutionary CoT Jailbreaks for LLMs
Jianan Li, Simeng Qin, Xiaojun Jia +5
Large Reasoning Models (LRMs) have demonstrated remarkable capabilities in reasoning and generation tasks and are increasingly deployed in real-world applications. However, their e…
Inverse Reinforcement Learning with Dynamic Reward Scaling for LLM Alignment
Ruoxi Cheng, Haoxuan Ma, Weixin Wang +7
Alignment is vital for safely deploying large language models (LLMs). Existing techniques are either reward-based (training a reward model on preference pairs and optimizing with r…