3 papers
cs.AI2026
Deactivating Refusal Triggers: Understanding and Mitigating Overrefusal in Safety Alignment
Zhiyu Xue, Zimo Qi, Guangliang Liu +2
Safety alignment aims to ensure that large language models (LLMs) refuse harmful requests by post-training on harmful queries paired with refusal answers. Although safety alignment…
cs.CL2025
Diagnosing the Performance Trade-off in Moral Alignment: A Case Study on Gender Stereotypes
Guangliang Liu, Bocheng Chen, Han Zi +2
Moral alignment has emerged as a widely adopted approach for regulating the behavior of pretrained language models (PLMs), typically through fine-tuning on curated datasets. Gender…
cs.CR2024
No Free Lunch for Defending Against Prefilling Attack by In-Context Learning
Zhiyu Xue, Guangliang Liu, Bocheng Chen +2
The security of Large Language Models (LLMs) has become an important research topic since the emergence of ChatGPT. Though there have been various effective methods to defend again…