15 papers
Between Safe Boundaries: Exploiting Temporal Consistency for Jailbreaking Text-To-Video Generation Models
Xingkai Peng, Jun Jiang, Jiayang Liu +2
Recently, text-to-video (T2V) models have been widely deployed, sparking growing concerns over their robustness against jailbreak attacks. Existing jailbreak methods, mostly adapte…
DNA: Dual-stage Native Attribution for Generated Image Source Tracing
Chao Wang, Kejiang Chen, Zijin Yang +4
The paper proposes DNA, a two‑stage framework that attributes generated images to their source models without additional training by first screening at the family level and then pi…
Membership Inference Attacks on Tokenizers of Large Language Models
Meng Tong, Yuntao Du, Kejiang Chen +2
Membership inference attacks (MIAs) are widely used to assess the privacy risks associated with machine learning models. However, when these attacks are applied to pre-trained larg…
Into the Gray Zone: Domain Contexts Can Blur LLM Safety Boundaries
Ki Sen Hung, Xi Yang, Chang Liu +7
A central goal of LLM alignment is to balance helpfulness with harmlessness, yet these objectives conflict when the same knowledge serves both legitimate and malicious purposes. Th…
InferDPT: Privacy-Preserving Inference for Closed-box Large Language Model
Meng Tong, Kejiang Chen, Jie Zhang +5
Large language models (LLMs), like ChatGPT, have greatly simplified text generation tasks. However, they have also raised concerns about privacy risks such as data leakage and unau…
AEDR: Training-Free AI-Generated Image Attribution via Autoencoder Double-Reconstruction
Chao Wang, Zijin Yang, Yaofei Wang +2
The rapid advancement of image-generation technologies has made it possible for anyone to create photorealistic images using generative models, raising significant security concern…