1 paper
Tiansheng Huang, Gautam Bhattacharya, Pratik Joshi +2
Safety aligned Large Language Models (LLMs) are vulnerable to harmful fine-tuning attacks -- a few harmful data mixed in the fine-tuning dataset can break the LLMs's safety alignme…