Showing cs.CRShow all
2 papers · 1 filter
cs.CR2025★ 1 cited
Uncovering and Aligning Anomalous Attention Heads to Defend Against NLP Backdoor Attacks
Haotian Jin, Yang Li, Haihui Fan +3
Backdoor attacks pose a serious threat to the security of large language models (LLMs), causing them to exhibit anomalous behavior under specific trigger conditions. The design of…
cs.CR2025
Fine-Tuning Jailbreaks under Highly Constrained Black-Box Settings: A Three-Pronged Approach
Xiangfang Li, Yu Wang, Bo Li
With the rapid advancement of large language models (LLMs), ensuring their safe use becomes increasingly critical. Fine-tuning is a widely used method for adapting models to downst…