3 papers
cs.CR2026
Rethinking Forgery Attacks on Semantic Watermarks in Black-Box Settings: A Geometric Distortion Perspective
Cheng-Yi Lee, Yichi Zhang, Yuchen Yang +2
Recent studies have shown that semantic watermarks, which embed information into the initial noise of latent diffusion models (LDMs), are vulnerable to black-box forgery attacks. H…
cs.LG2025
Safety Alignment Depth in Large Language Models: A Markov Chain Perspective
Ching-Chia Kao, Chia-Mu Yu, Chun-Shien Lu +1
Large Language Models (LLMs) are increasingly adopted in high-stakes scenarios, yet their safety mechanisms often remain fragile. Simple jailbreak prompts or even benign fine-tunin…
cs.LG2024
The Great Contradiction Showdown: How Jailbreak and Stealth Wrestle in Vision-Language Models?
Ching-Chia Kao, Chia-Mu Yu, Chun-Shien Lu +1
Vision-Language Models (VLMs) have achieved remarkable performance on a variety of tasks, yet they remain vulnerable to jailbreak attacks that compromise safety and reliability. In…