3 papers
cs.LG2025
SafeLLM: Unlearning Harmful Outputs from Large Language Models against Jailbreak Attacks
Xiangman Li, Xiaodong Wu, Qi Li +2
Jailbreak attacks pose a serious threat to the safety of Large Language Models (LLMs) by crafting adversarial prompts that bypass alignment mechanisms, causing the models to produc…
cs.CR2025
Robustness Assessment and Enhancement of Text Watermarking for Google's SynthID
Xia Han, Qi Li, Jianbing Ni +1
Recent advances in LLM watermarking methods such as SynthID-Text by Google DeepMind offer promising solutions for tracing the provenance of AI-generated text. However, our robustne…
cs.CR2025
SecureT2I: No More Unauthorized Manipulation on AI Generated Images from Prompts
Xiaodong Wu, Xiangman Li, Qi Li +2
Text-guided image manipulation with diffusion models enables flexible and precise editing based on prompts, but raises ethical and copyright concerns due to potential unauthorized…