7 papers
Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates
Weiyi He, Yuping Lin, Jiliang Tang +1
Adversarial training is one of the most effective defenses against adversarial attacks, yet the computational cost remains prohibitive at modern scales, especially for large langua…
Six-CD: Benchmarking Concept Removals for Benign Text-to-image Diffusion Models
Jie Ren, Kangrui Chen, Yingqian Cui +5
Text-to-image (T2I) diffusion models have shown exceptional capabilities in generating images that closely correspond to textual prompts. However, the advancement of T2I diffusion…
Unveiling and Mitigating Memorization in Text-to-image Diffusion Models through Cross Attention
Jie Ren, Yaxin Li, Shenglai Zeng +4
Recent advancements in text-to-image diffusion models have demonstrated their remarkable capability to generate high-quality images from textual prompts. However, increasing resear…
Data Poisoning for In-context Learning
Pengfei He, Han Xu, Yue Xing +3
In the domain of large language models (LLMs), in-context learning (ICL) has been recognized for its innovative ability to adapt to new tasks, relying on examples rather than retra…
Sharpness-Aware Data Poisoning Attack
Pengfei He, Han Xu, Jie Ren +4
Recent research has highlighted the vulnerability of Deep Neural Networks (DNNs) against data poisoning attacks. These attacks aim to inject poisoning samples into the models' trai…
Multi-Faceted Studies on Data Poisoning can Advance LLM Development
Pengfei He, Yue Xing, Han Xu +2
The lifecycle of large language models (LLMs) is far more complex than that of traditional machine learning models, involving multiple training stages, diverse data sources, and va…