21 citations · 21 across the 2 of their papers we have counts for
1 paper · 1 filter
Tiansheng Huang, Gautam Bhattacharya, Pratik Joshi +2
Safety aligned Large Language Models (LLMs) are vulnerable to harmful fine-tuning attacks -- a few harmful data mixed in the fine-tuning dataset can break the LLMs's safety alignme…