From the 1 of 5 linked papers with an AI index.
5 papers
Crossing the Margin Cliff: Toward Relearn-Robust LLM Unlearning via Margin Calibration
Xiangyu Yin, Jiaxu Liu, Zhen Chen +1
The paper identifies a systematic “margin cliff” that makes large language model unlearning vulnerable to relearn attacks and proposes Margin Calibration, a plug‑in method that add…
Fragile by Design: On the Limits of Adversarial Defenses in Personalized Generation
Zhen Chen, Yi Zhang, Xiangyu Yin +4
Personalized AI applications such as DreamBooth enable the generation of customized content from user images, but also raise significant privacy concerns, particularly the risk of…
Trustworthy Text-to-Image Diffusion Models: A Timely and Focused Survey
Yi Zhang, Zhen Chen, Chih-Hong Cheng +6
Text-to-Image (T2I) Diffusion Models (DMs) have garnered widespread attention for their impressive advancements in image generation. However, their growing popularity has raised et…
TAIJI: Textual Anchoring for Immunizing Jailbreak Images in Vision Language Models
Xiangyu Yin, Yi Qi, Jinwei Hu +5
Vision Language Models (VLMs) have demonstrated impressive inference capabilities, but remain vulnerable to jailbreak attacks that can induce harmful or unethical responses. Existi…
CeTAD: Towards Certified Toxicity-Aware Distance in Vision Language Models
Xiangyu Yin, Jiaxu Liu, Zhen Chen +4
Recent advances in large vision-language models (VLMs) have demonstrated remarkable success across a wide range of visual understanding tasks. However, the robustness of these mode…