10 papers
InferDPT: Privacy-Preserving Inference for Closed-box Large Language Model
Meng Tong, Kejiang Chen, Jie Zhang +5
Large language models (LLMs), like ChatGPT, have greatly simplified text generation tasks. However, they have also raised concerns about privacy risks such as data leakage and unau…
Character as a Latent Variable in Large Language Models: A Mechanistic Account of Emergent Misalignment and Conditional Safety Failures
Yanghao Su, Wenbo Zhou, Tianwei Zhang +4
Emergent Misalignment refers to a failure mode in which fine-tuning large language models (LLMs) on narrowly scoped data induces broadly misaligned behavior. Prior explanations mai…
EditMark: Watermarking Large Language Models based on Model Editing
Shuai Li, Kejiang Chen, Jun Jiang +5
Large Language Models (LLMs) have demonstrated remarkable capabilities, but their training requires extensive data and computational resources, rendering them valuable digital asse…
SafeGuider: Robust and Practical Content Safety Control for Text-to-Image Models
Peigui Qi, Kunsheng Tang, Wenbo Zhou +5
Text-to-image models have shown remarkable capabilities in generating high-quality images from natural language descriptions. However, these models are highly vulnerable to adversa…
On the Vulnerability of Text Sanitization
Meng Tong, Kejiang Chen, Xiaojian Yuan +4
Text sanitization, which employs differential privacy to replace sensitive tokens with new ones, represents a significant technique for privacy protection. Typically, its performan…
Clean Image May be Dangerous: Data Poisoning Attacks Against Deep Hashing
Shuai Li, Jie Zhang, Yuang Qi +4
Large-scale image retrieval using deep hashing has become increasingly popular due to the exponential growth of image data and the remarkable feature extraction capabilities of dee…