3 papers
cs.LG2025
Safety Alignment Depth in Large Language Models: A Markov Chain Perspective
Ching-Chia Kao, Chia-Mu Yu, Chun-Shien Lu +1
Large Language Models (LLMs) are increasingly adopted in high-stakes scenarios, yet their safety mechanisms often remain fragile. Simple jailbreak prompts or even benign fine-tunin…
cs.LG2024
The Great Contradiction Showdown: How Jailbreak and Stealth Wrestle in Vision-Language Models?
Ching-Chia Kao, Chia-Mu Yu, Chun-Shien Lu +1
Vision-Language Models (VLMs) have achieved remarkable performance on a variety of tasks, yet they remain vulnerable to jailbreak attacks that compromise safety and reliability. In…
cs.CV2024
Defending Against Repetitive Backdoor Attacks on Semi-supervised Learning through Lens of Rate-Distortion-Perception Trade-off
Cheng-Yi Lee, Ching-Chia Kao, Cheng-Han Yeh +3
Semi-supervised learning (SSL) has achieved remarkable performance with a small fraction of labeled data by leveraging vast amounts of unlabeled data from the Internet. However, th…